18 sep
|
Lilt
|
Miguel Hidalgo
18 sep
Lilt
Miguel Hidalgo
About The OpportunityWe are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.
We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches.
Note this is a remote, freelance opportunity
What You’ll Deliver
- Task Engineering: Evaluating Coding Agents.
- Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling.
- Prompting & Translation: finding failure points where AI does not work, in your native language
- Implementation & Verification:
Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary).
- Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).
- Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.
Qualifications
- Experience: 5+ years of industry experience in software engineering.
- Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities.
- Language: Native or near-native fluency, with a deep understa
📌 AI Benchmark Engineer | Native Language Specialist (Miguel Hidalgo)
🏢 Lilt
📍 Miguel Hidalgo