Four ways to deploy an agent, keyed to the one thing that actually varies — the window between the work arriving and the answer being due. Each lane runs its flow dot at its own speed, so batch crawls where request snaps, and each ends on the failure that comes attached to its shape.
the same agent, four answer windows — schematic, not measured
Nothing here is a preference. The window between the work arriving and the answer being due is set by the product, and it decides the rest: whether a queue can absorb a spike, whether replicas need warm models, whether there is a network at all. Pick the shape that fits the window, then live with the failure that comes attached to it.