"Small" was a poor choice of words here, "low compute budget" is more what I'm getting at.
In voice interactions, ttfat is actually relatively important. If you look at models with a <1s ttfat you eliminate almost every reasoning model, less some of the diffusion models and more obscure ddtree/dflash like speculative decoding implementations.
GPT-Live, which is coming to the API soon, responds instantly while reasoning in the background. So it can say "Hold on, I'll look that up for you" and continue to respond to the user conversationally while running an asynchronous reasoning task in the background.
It's not in the API yet, but it should be in the coming weeks. You'll see an enormous improvement compared to GPT-4.1.
What? GPT-4.1 was not a small model! And why wouldn't you use reasoning?
You're of course going to see poor results when you restrict yourself to small non-reasoning models, but why would you?