How we cut time to first byte by 94% for long prompts

LiteLLM was counting every token in a long conversation to answer a yes-or-no routing question.
Removing that unnecessary work took median time to first byte from 553 ms to 35 ms in our local 440k-token benchmark: 94% lower.






