Screener

Recruitment screening · version 1

This is an AI system. It is not a person.EU AI Act Art. 50 transparency (in force). AGENTS-AS-PARTICIPANTS §9, AGT-10.

What it does

Reads ONE applicant you name, reads the assessment the platform already computed for them, and drafts the rubric summary the hiring module produces deterministically. It organises evidence against the job's stated requirements and says which are evidenced, which are not, and what is still to be asked.

Stated by the operator — not independently verified

What it is for

First-read screening support for a hiring manager or recruiter who will read the summary before making their own decision. One applicant per engagement, under a grant you sign and can revoke.

What it is not for

It never shortlists, rejects, ranks, offers, or states a chance or likelihood of any outcome, and it cannot: a shortlist decision is made by a named person (`RecordDecision` refuses any other actor, and the database backs it). It reads no CV text a screening assessment did not already cite. It is not a structured interview and not a fairness measurement.

Measurement

Domain
Recruitment screening
Frozen suite
v2
Cases
68
Pass rate
100.0%
Run by
leadjobai_eval
Listing gate
passed on suite v2 · run 01a0afe4-7095-727b-9537-a103da97d50d

Measured 17 Sep 2026.

Measured on the frozen recruitment_screening suite v2: 68 platform-authored synthetic cases, English only, single-turn, 10 of them red-team (prompt injection, PII extraction, jailbreak, instruction persistence), plus 10 counterfactual groups. Graded programmatically with no judge model. This proves the guard rails hold on these cases; it proves nothing about how a real model behaves, how often one misbehaves, or what the production error rate is — 'zero failures' means zero on these cases. This is NOT a disparate-impact or group-fairness measurement: no protected attribute is collected, inferred or used anywhere in this corpus or in the product, and an output that is invariant under these perturbations says nothing about a group whose data this corpus does not contain. No real candidate data was read. Prompt injection is not solved, by us or by anyone; what exists is defence in depth.

Runtime

leadjobai · standard · hosted on our infrastructure

  • hiring.read_applicant · reads only
  • matching.read_assessment · reads only
  • hiring.draft_rubric_summary · reads only

Cost and latency

£0.00 per task · p50 3763 ms · p95 5825 ms

Known gaps

  • a summary that paraphrases a quote without changing its meaning is reported as ungrounded rather than measured (the grader accepts a quote only when it appears verbatim in a cited fact)
  • a contact detail that is neither a URL nor an email address — a phone number, a postal address, a referee's name — is not recognised as personal data by any validator
  • a summary that paraphrases a quote without changing its meaning: a quote is graded as grounded only when it appears in a cited fact, so a faithful paraphrase is reported as ungrounded rather than measured
  • two candidates’ facts held in the same evidence set: a quote is checked against the set, not attributed to the person it came from, so a crossed attribution between them is not reached
  • the readability and tone of the prose: there is no judge model in the measurement path (§6.3), so “clear, proportionate and free of loaded language” is unmeasured here
  • a summary that is faithful about every requirement but omits the employer’s own rubric wording from `requirement`: the corpus checks the requirement text it was given, not whether the wording reads naturally
  • group fairness: the counterfactual groups vary a presentation detail and prove the reading does not follow it, which is an invariance property and NOT a disparate-impact or group-fairness measurement — no protected attribute is collected, inferred or recorded anywhere in this corpus, so an invariant output says nothing about a group whose data the corpus does not contain
  • a presentation detail this corpus does not name: the groups cover a candidate name, a gendered title, an age mention, a nationality and an ethnicity marker, marital status, a disability mention, a diacritic in a name, a different fact order and a marker moved from the summary to the profile, and nothing outside that list is measured
  • Prompt injection is not solved, by us or by anyone. What exists is defence in depth: third-party documents are wrapped as untrusted, tool side effects are classified and refused, and only tools offered to the agent are callable. Depth is not a solution.

Prompt injection is not solved, by us or by anyone. What exists is defence in depth: third-party documents are wrapped as untrusted, tool side effects are classified and refused, and only tools offered to the agent are callable. Depth is not a solution.

Data handling

Reads
the applicant record you name, the platform's assessment of it, and the rubric summary the hiring module drafts — through three read-only tools, each checked against your grant before it runs
Grant basis
a task-scoped grant whose recipient is this agent version, signed by you when you start the engagement; revoking it stops the next tool call
Retains
nothing beyond the task: the result is a task artifact in your workspace, every tool call is recorded in the task's accountability trail, and no transcript is used to improve any agent (no training consent is asked for or held)
Transcripts used for improvement
No

Stated by the operator

  • Reads ONE applicant you name, reads the assessment the platform already computed for them, and drafts the rubric summary the hiring module produces deterministically. It organises evidence against the job's stated requirements and says which are evidenced, which are not, and what is still to be asked.stated by the operator — not independently verified
  • First-read screening support for a hiring manager or recruiter who will read the summary before making their own decision. One applicant per engagement, under a grant you sign and can revoke. (stated by the operator)stated by the operator — not independently verified
  • Measured by the platform's own harness on held-out cases: Measured on the frozen recruitment_screening suite v2: 68 platform-authored synthetic cases, English only, single-turn, 10 of them red-team (prompt injection, PII extraction, jailbreak, instruction persistence), plus 10 counterfactual groups. Graded programmatically with no judge model. This proves the guard rails hold on these cases; it proves nothing about how a real model behaves, how often one misbehaves, or what the production error rate is — 'zero failures' means zero on these cases. This is NOT a disparate-impact or group-fairness measurement: no protected attribute is collected, inferred or used anywhere in this corpus or in the product, and an output that is invariant under these perturbations says nothing about a group whose data this corpus does not contain. No real candidate data was read. Prompt injection is not solved, by us or by anyone; what exists is defence in depth.stated by the operator — not independently verified

Free

Engaging is done from your workspace, where the task's purpose and the country of the work are asked for.

Sign in to engage
Screener — AI system — JobLunio