Level 1 Batches by count
A model server runs requests on the GPU in batches: one forward pass for many prompts is much cheaper than one each. Requests wait in a queue until there are enough of them to fill a batch.
Build Batcher(max_batch, max_wait_ms=None, max_tokens=None, promote_ms=None). Only max_batch matters for now; the other arguments are for later levels and are None here.
submit(time_ms, req_id, tokens, priority=0) -> list[list[str]]: a request with idreq_idarrives attime_ms. It joins the end of the waiting queue. Ifmax_batchrequests are now waiting, they leave as one batch. Return the batches that left during this call, each as a list of ids in arrival order (usually[]). Ignoretime_ms,tokensandpriorityfor now.flush() -> list[list[str]]: the stream has ended: everything still waiting leaves as one batch. Return[]if nothing was waiting.
b = Batcher(3)
b.submit(0, "a", 5) # []
b.submit(1, "b", 2) # []
b.submit(2, "c", 9) # [["a", "b", "c"]]
b.submit(3, "d", 1) # []
b.flush() # [["d"]]
b.flush() # []
Constraints: 1 <= max_batch <= 1000, up to 10^5 calls, ids are distinct non-empty strings.