AI Benchmark Engineer - Native Language Specialist | Korean 48 applicants
The engineer will design, build, and validate rigorous evaluation benchmarks for large language models focusing on multilingual software challenges within terminal workflows. This involves creating realistic task environments, identifying model failure points in native languages, and developing reliable verification scripts.