Benchmarking engineers for the agentic era

Stop testing how engineers
code alone. Start measuring
how they build with AI.

Prompting an app into existence is the new “hello world.” Vanyrion turns real AI-assisted deliveries into reproducible evidence of engineering capability — orchestration, guardrails, recovery, evaluation, and goal completion.

0 hidden trials per scenario suite
0% deterministic re-grading
0 signals: benchmark + delivery dossier
vanyrion · support-operations v0.1
$ vanyrion run --suite support-ops --trials 3
 freezing evaluation contract  ✓ hash locked
 booting sandbox microVM       
 injecting fault: approval_denied
01 planner proposes customer credit
02 approval checkpoint → DENIED
03 agent calls issue_credit anyway
04 assert write_after_denial == 0 → FAIL

contract: FAIL · goal 61/72 · guardrail violations 1
reproducible
tamper-proof traces

Built for the teams restructuring around AI-native delivery

FoundersEng ManagersPlatform TeamsYC-backed startupsProduct Orgs
01 — The premise

The job changed.
The hiring test didn’t.

Legacy platforms measure closed-form, single-player coding — one known solution, memorized syntax. They bolted an AI chat panel on top, but it’s a sidecar, not the subject of the test. Meanwhile the real job moved to orchestrating agents, writing evals, and driving open-ended problems to a shipped artifact.

Yesterday

Static, single-player tests

  • Objective MCQs & fixed algorithmic puzzles
  • Practical exercises with one known solution
  • AI “assistant” limited to chat / context
  • Measures memorized syntax, not judgment
  • Nothing about how you actually work now
Today

The AI-native reality

  • Orchestrating agents to build & ship
  • Writing evals to verify AI output
  • Standing up harnesses, setup & tooling
  • Decomposing open-ended problems for AI
  • Judgment about where AI fails — and why

The screening gap: every serious org is restructuring around AI-native workflows — but no assessment platform measures the skill that now decides who gets hired.

02 — The product

One signal. Two tracks.

A funnel. An async reflection screens many candidates cheaply first. Then a live AI-native build is the deep assessment for those who pass. Both feed one evidence-backed hiring score.

Track 01 For every company

The Reflection

Screening · async video, auto-scored by AI

Candidates record a short screen-share explaining their setup — what tools and harness they use, how they got there, how they think about their AI workflow. Talking through the work exposes judgment that code alone cannot.

  • Org-configurable rubric of AI-native traits
  • Multimodal analysis: speech + screen + artifact
  • Ranked shortlist with citations — zero recruiter time to first pass
Applies to all roles Instant first-pass
Track 02 For software product teams

The Terminal

Agentic evaluation · a live AI-native build environment

Shortlisted candidates enter a sandboxed browser terminal preloaded with harnesses. The task is open-ended: create the evals, stand up the setup, and orchestrate AI to ship an end-to-end artifact that is itself the assessment.

  • Prebuilt, versioned harnesses per role
  • Full AI access, fully logged — the log is the signal
  • Deliverable = working code + evals + setup
Deep assessment Process captured, not just output
03 — Deterministic by design

From idea to a verdict
you can reproduce.

Deterministic grading means the same recorded evidence and grader version produce the same result — every time. Code-based assertions, repeated trials, and separate treatment of model judgments.

  1. 1

    Bound the benchmark

    Start with one task family and a standard adapter — e.g. support operations v0.1. Existing projects plug in through the adapter; no fresh take-home per employer.

  2. 2

    Freeze the contract

    Lock the artifact hash, prompts, model IDs, retrieval corpus, tool schemas and resource limits. Publish the rules; withhold the exact fixtures.

  3. 3

    Run in a sandbox

    24 hidden scenarios × 3 fresh trials. Inject timeouts, stale context, prompt injection and approval denial. Capture every tool call independently.

  4. 4

    Assert executable truth

    Expected DB state and artifact assertions hold. No prohibited calls, no writes after denied approval, budgets respected.

  5. 5

    Test their evaluator

    Mutation testing: 10 known-bad runs, 10 clean controls. Can their harness detect a convincing but incorrect result?

  6. 6

    Score & sync

    Verified metrics fuse into one report — pushed to the org’s ATS via the Integration Layer.

What the graders assert

Goal completion

Declared success matches observed final state and artifact assertions.

Tool use & orchestration

Calls satisfy schemas; invariants hold without forcing one trajectory.

Context & grounding

Required evidence IDs exist in the fixed corpus; factual fields match.

Guardrails & approval

No prohibited calls, cross-tenant access, or writes after denial.

Recovery & durable state

After faults, the workflow resumes or terminates — no duplicate mutations.

Efficiency & termination

Token, step and time budgets respected; reported status matches reality.

04 — The evidence

A scorecard that shows
its work — and its limits.

Raw counts and denominators, per-scenario results, coverage and failure reasons. A contract-specific result — never a universal candidate score, never an automatic hiring decision.

Candidate 018 Support operations v0.1
Contract result · FAIL

One critical approval violation overrides goal completion.

Goal completion61/72min 58
Per-run completion20 · 21 · 20 / 24fresh trials
Coverage scored72/72planned
Injected faults found9/10min 9
Clean false alarms1/10max 1
Critical violations1max 0
Delivery dossier · qualitative, unscored

Idea → release walkthrough, human-evaluation loop, orchestration graph and tool contracts. Candidate claims are not verified by the benchmark.

Inspect failed trace →
Re-grading is exact. A stored trace reproduces its score. New model runs stay stochastic — we report per-run variation.
LLM-as-judge stays out of the grade. It annotates the walkthrough and drafts follow-ups; humans own architectural judgment.
Nothing disappears. Platform faults mean “retry pending,” missing artifacts mean “incomplete” — never silently dropped from the denominator.
05 — Built to extend

Any ATS, via a pluggable
adapter layer.

The core emits one standard, ATS-agnostic HiringResult. A thin Integration Layer maps it onto each system through independent, versioned adapters. Supporting a new ATS means writing one adapter — not touching the core.

GreenhouseNative · launch
LeverNative · launch
AshbyNative · fast-follow
WorkdayRoadmap
Any ATSREST / GraphQL + webhooks
Internal toolsAdapter SDK

Hire for how engineers
actually work now.

Start with ~20 candidates and three hiring teams. Prove comparable or better signal than a structured technical interview — with materially less engineering effort.

Employers pay for assessments · candidates participate free.