Signal & Eval Engineer
Mid-level, Emergences Labs, AI Evaluation & Workforce Intelligence, San Francisco, CA (remote-friendly), Remote
About the team
The engineering team is small, ships daily, and uses Claude and Cursor as primary collaborators. Founders code. AI-generated diffs get more review scrutiny, not less. No ticket queues, no separate DevOps layer, no permission culture — if you see something broken, you fix it.
About the company
Emergences Labs builds AI Evaluation & Workforce Intelligence Infrastructure. The NeoHuman platform captures what candidates actually do during work sessions — desktop behavior, tool use, reasoning traces — and turns it into defensible hiring evidence for enterprise customers. Three service lines: Assessment (core revenue), Training, and Data. SF Bay Area, distributed team.
About the role
NeoHuman turns real candidate behavior into hiring evidence. The LLM scoring and signal extraction pipeline is the system that decides whether that evidence is exact or inferred — a distinction that defines the quality of every hiring decision we deliver to enterprise customers. You will own this pipeline end-to-end: the prompts, the eval loop, and the quality bar that keeps it honest in production.
What you'll do
- Own the exact vs inferred classification pipeline — design and maintain the prompts, schemas, and decision logic that tag candidate behavior as exact evidence or inferred signal; the quality of every hiring-manager report flows through this.
- Run the eval loop — build and maintain regression test suites that catch scoring drift before it reaches production; decide when a prompt or schema change improves signal fidelity and when it quietly degrades it.
- Debug production LLM failures — when the pipeline produces a wrong output, diagnose the root cause (wrong context window, schema ambiguity, sensor artifact, model drift) and fix it; not re-prompt and hope.
- Iterate on extraction prompts with Claude as a collaborator — structure problems so the model executes reliably, validate its output rigorously, and own the final call on what behavioral signal counts.
- Deliver scoring templates for new assessment configurations — translate rubric designs into pipeline-ready templates; review AI-generated scaffolding for accuracy before each design-partner run.
- Instrument the pipeline — track cost-per-session, latency, and evidence quality metrics; surface regressions early and close the loop before they reach customers.
Human–AI ways of working
You work with Claude and Cursor as primary collaborators, not shortcuts. Claude drafts prompt variants, generates regression test cases, and scaffolds scoring templates; you structure the problems, validate the outputs, and own every call on what counts as exact evidence. The distinction matters here more than almost anywhere else: AI proposes, you decide. Being AI-fluent in this role means knowing when the model is confidently wrong — and being able to explain why to a hiring manager whose decision rests on it.
Required qualifications
- Shipped a production LLM pipeline — not a prototype — that processed real user data and held up under edge cases; owned it through incidents, not just launch.
- Can debug a bad LLM output and explain the root cause: wrong context window, schema ambiguity, sensor artifact, model drift — not just re-prompt.
- Has owned an eval loop: built regression tests for a prompt-based pipeline, tracked quality drift over time, and made the call on when a change degrades signal fidelity.
- Proficient in TypeScript and/or Python; comfortable with the Anthropic Claude API and a PostgreSQL/Supabase-backed stack.
Preferred qualifications
- Experience with multi-modal or behavioral signal extraction — text combined with UI events, keyboard/mouse traces, or other fused, noisy sensor data.
- Has worked in a consequential classification context (hiring, compliance, legal, or medical) where a misclassification has real downstream stakes.
- Familiarity with human annotation workflows and inter-annotator agreement metrics.
Reports to
CTO & Cofounder
Employment type
Full-time
Compensation
$150,000–$180,000 base + Series B equity
How to apply
Send a short note to jack@emergences.ai describing a production LLM pipeline you have owned — what it did, what broke, and how you diagnosed it. No cover letter. No recruiter screen.