AI Benchmark Engineer | Native Language Specialist - Arabic (UAE) - Remote
LILT (Production)· United Arab Emirates (remote)
remote seen live 16h ago
About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects,
Apply on LILT (Production)'s site
This link goes straight to the employer's ashby page. Posted 10 days ago.We last confirmed it was open 16h ago.