Level 1 Pass rates
Before a new model ships, it runs through an evaluation suite: hundreds of test cases, each run several times because the model's answers vary. Results stream in from many workers, and the team wants live numbers. Over four levels you'll build EvalHarness.
EvalHarness(): starts empty.record(case_id, tags, passed, run_index) -> None: one result. Casecase_id(a string) was run for therun_index-th time and passed or failed.tagslists the case's distinct tags, like["math", "long"]; it may be empty, and a case always arrives with the same tags. If this(case_id, run_index)already has a result, the new one replaces it (the run was retried).pass_rate() -> float: passed results divided by all results.0.0if there are none.tag_pass_rate(tag) -> float: the same, counting only results of cases that havetag.0.0if no case has it.
In this level, whenever a rate is asked for, every case recorded so far has the same number of runs. Answers within 10^-6 of the expected value are accepted.
h = EvalHarness()
h.record("add", ["math"], True, 0)
h.record("sub", ["math"], False, 0)
h.record("greet", ["chat"], True, 0)
h.pass_rate() # 0.6667 (2 of 3)
h.tag_pass_rate("math") # 0.5
h.record("add", ["math"], True, 1)
h.record("sub", ["math"], True, 1)
h.record("greet", ["chat"], True, 1)
h.pass_rate() # 0.8333 (5 of 6)
h.record("sub", ["math"], True, 0) # run 0 of "sub" was retried and passed
h.tag_pass_rate("math") # 1.0
h.tag_pass_rate("code") # 0.0
Up to 200,000 calls; run_index is in 0..10^6. ⭐ Bonus: a speed test earns a star if every call is O(1) (apart from the number of tags).