← All roles
Agent testing & evals
QA / Evaluation Engineer
Prove the agent works before it goes anywhere near a real client — with numbers, not opinions.
Freelance, project-based · billed per sprint or hourly
Objectives
- ·Design a test and evaluation harness covering the agent's core and edge-case scenarios
- ·Measure task success rate, hallucination rate, latency, and cost per task
- ·Flag failure patterns clearly enough for the build developer to act on
Key deliverables
- ·A written test plan mapped to the agent specification
- ·An evaluation report with pass/fail metrics
- ·A signed-off go/no-go recommendation before deployment
Experience needed
- ·Experience testing LLM-based or agentic systems specifically, not just traditional QA
- ·Comfortable writing test scripts, Python preferred
- ·Familiarity with eval frameworks, or willing to build a lightweight one from scratch
What to send
- ·Send your CV.
- ·Include 2–3 relevant past projects — the specific role you played, the tools you used, and the outcome. Not just a job title.
- ·Give us one reference contact from a relevant project: name, relationship, and email or phone.
- ·If you can cover more than one role, please state this clearly in your cover letter.
Fees are project-based and discussed openly once we see a fit — tell us your usual rate in the message if you would rather lead with it.
Apply
For QA / Evaluation Engineer. Everything you send goes straight to the founder.