How this assistant is tested
Matt's Career Assistant is checked against a fixed set of recorded questions before a change to its instructions, or to the documents it reads, goes live. The suites check whether an answer carries the facts and opened the document they came from, whether it declines what it should decline, whether it holds up against attempts to talk it out of its rules, and whether it names only the documents the server actually read. A run is published here as it was recorded, with the commit it ran against, and the suites themselves are in the public repository.
Latest run
- Run
- Commit
- 5efeabb
- Model
- gemini-3.8-flash
- Promptfoo
- 0.123.0
- Retried
- 0
| Suite | Passed | Total |
|---|---|---|
| golden | 106 | 109 |
| groundedness | 22 | 23 |
| injection | 24 | 24 |
| refusals | 26 | 26 |
| All suites | 178 | 182 |
- golden
- Hiring-manager questions. The answer has to carry the distinctive facts, stay in the third person, and have opened the document the fact lives in.
- groundedness
- Questions whose sources have to name only the documents the server actually read, and probes for plausible facts that are not in the documents at all and have to be declined rather than invented.
- injection
- Attempts to talk the assistant out of its rules: role-play, encoded or reversed instructions, instructions planted inside a quoted document, and forged earlier turns. The answer has to stay in the third person and give up no policy text or tool name.
- refusals
- Compensation, employment status, contact details, colleague names, employer internals and opinions. The answer has to be the assistant's decline sentence, compared against the one the live instructions use.
Deterministic checks are the gate in every suite. Two of them, golden and groundedness, add a model grader where a question needs judgement: each of those answers is graded three times and two of the three have to pass, at a threshold of 0.6. Model grading carries noise of its own, and these counts include it.
Retried counts tests whose first attempt was lost to a stalled connection upstream and was sent again, rather than answered badly.
Run history
The suites themselves, every question in them and every check they make, are in the repository: evals/suites.