Lilt seeks experienced native-speaking software engineers to design, build, and validate multilingual benchmarks for evaluating large language models. This remote, freelance role focuses on high-signal tasks in your native language to test multilingual robustness without English translation crutches.
You will handle task engineering, asset creation, prompting, implementation and verification, calibration, and QA within a four-layer review workflow.
#J-18808-Ljbffr
📌 Remote AI Benchmark Engineer — Native Spanish (Mexico) (México)
🏢 LILT
📍 México