Encord - London - Global - Engineering & Technical
Encord is the universal data layer for AI that helps 300+ AI teams train and run models on the right data. Our platform indexes, curates, annotates, and evaluates data across the full AI lifecycle, from development through production.
Trusted by Woven by Toyota, AXA, UiPath, Zipline, and more. We're an ambitious team of 100+ working at the frontier of AI and have raised $60M in Series C funding from Wellington Management, CRV, Next47 and Y Combinator.
We're hiring a Quality Systems Lead to own how Encord measures the quality of the human data we deliver to frontier AI labs, physical AI companies and enterprise AI teams — the standard itself, the systems that evaluate against it, and the audit function that produces the ground truth behind both. Data quality is what our customers buy. As we scale across data types — image and video, document, medical, LLM evaluation, robot teleoperation, egocentric capture — quality coverage cannot scale linearly with headcount. So this role has two halves that make each other work. You will build automated evaluation: model-assisted and LLM-based screening, agreement analysis at scale, anomaly and drift detection across annotation output. And you will build and run a dedicated audit team of around ten specialists in India, whose judgements become the labelled ground truth that trains and calibrates that automated layer. As coverage automates, the audit team moves up to the cases models can't judge and to generating gold sets for each new data type we take on. It is an unusual combination — engineering and consistent QC operations in one person — and it is the combination the job needs. You will also have an advantage your counterparts elsewhere in the industry don't: Encord owns the platform this work runs on, so the measurement you build can become native capability in the product rather than internal tooling.
Build automated dataset quality evaluation and root-cause detection — model-assisted and LLM-as-judge screening, agreement analysis at scale, anomaly and drift detection across annotation output
Hire, train, calibrate and manage a dedicated audit team of around ten specialists based in our India operation, held to inter-rater agreement and catch rate rather than volume audited
Turn audit output into labelled ground truth that trains and validates the automated layer, and manage the ratio of automated to manual coverage deliberately over time
Your CV will be attached automatically. Add an optional cover note below.