~/problems / AI infrastructure

Dynamic request batcher

On a phone? Coding is easier on a laptop: email this problem to yourself . Meanwhile: fight a boss.

medium assessment 4 levels ~55 min

Level 1 Batches by count

A model server runs requests on the GPU in batches: one forward pass for many prompts is much cheaper than one each. Requests wait in a queue until there are enough of them to fill a batch.

Build Batcher(max_batch, max_wait_ms=None, max_tokens=None, promote_ms=None). Only max_batch matters for now; the other arguments are for later levels and are None here.

  • submit(time_ms, req_id, tokens, priority=0) -> list[list[str]]: a request with id req_id arrives at time_ms. It joins the end of the waiting queue. If max_batch requests are now waiting, they leave as one batch. Return the batches that left during this call, each as a list of ids in arrival order (usually []). Ignore time_ms, tokens and priority for now.
  • flush() -> list[list[str]]: the stream has ended: everything still waiting leaves as one batch. Return [] if nothing was waiting.
b = Batcher(3)
b.submit(0, "a", 5)   # []
b.submit(1, "b", 2)   # []
b.submit(2, "c", 9)   # [["a", "b", "c"]]
b.submit(3, "d", 1)   # []
b.flush()             # [["d"]]
b.flush()             # []

Constraints: 1 <= max_batch <= 1000, up to 10^5 calls, ids are distinct non-empty strings.

Level 2 unlocks when level 1 passes.

Level 3 unlocks when level 2 passes.

Level 4 unlocks when level 3 passes.

Topic: AI infrastructure. The plumbing around models: request batching, streaming responses, prompt caches, sampling, token limits and eval harnesses.

0:00
Ctrl ' run · Ctrl ↵ submit
esc