~/problems / AI infrastructure

Eval harness

On a phone? Coding is easier on a laptop: email this problem to yourself . Meanwhile: fight a boss.

medium assessment 4 levels ~45 min

Level 1 Pass rates

Before a new model ships, it runs through an evaluation suite: hundreds of test cases, each run several times because the model's answers vary. Results stream in from many workers, and the team wants live numbers. Over four levels you'll build EvalHarness.

  • EvalHarness(): starts empty.
  • record(case_id, tags, passed, run_index) -> None: one result. Case case_id (a string) was run for the run_index-th time and passed or failed. tags lists the case's distinct tags, like ["math", "long"]; it may be empty, and a case always arrives with the same tags. If this (case_id, run_index) already has a result, the new one replaces it (the run was retried).
  • pass_rate() -> float: passed results divided by all results. 0.0 if there are none.
  • tag_pass_rate(tag) -> float: the same, counting only results of cases that have tag. 0.0 if no case has it.

In this level, whenever a rate is asked for, every case recorded so far has the same number of runs. Answers within 10^-6 of the expected value are accepted.

h = EvalHarness()
h.record("add", ["math"], True, 0)
h.record("sub", ["math"], False, 0)
h.record("greet", ["chat"], True, 0)
h.pass_rate()               # 0.6667  (2 of 3)
h.tag_pass_rate("math")     # 0.5
h.record("add", ["math"], True, 1)
h.record("sub", ["math"], True, 1)
h.record("greet", ["chat"], True, 1)
h.pass_rate()               # 0.8333  (5 of 6)
h.record("sub", ["math"], True, 0)   # run 0 of "sub" was retried and passed
h.tag_pass_rate("math")     # 1.0
h.tag_pass_rate("code")     # 0.0

Up to 200,000 calls; run_index is in 0..10^6. ⭐ Bonus: a speed test earns a star if every call is O(1) (apart from the number of tags).

Level 2 unlocks when level 1 passes.

Level 3 unlocks when level 2 passes.

Level 4 unlocks when level 3 passes.

Topic: AI infrastructure. The plumbing around models: request batching, streaming responses, prompt caches, sampling, token limits and eval harnesses.

0:00
Ctrl ' run · Ctrl ↵ submit
esc