I've been thinking a lot about the interfaces we use to interact with large language models — and how LLMs interact with us.
Right now, the default input is text. You type something in, the model responds with more text. That works, but it's limiting. Especially when you consider the full spectrum of human communication.
We've seen a lot of progress on the output side: LLMs can now generate text, code, images, videos, Figma designs, and more. But we need to refocus on the input side.
Text shouldn't be the only way to communicate with models. Audio should become a default input interface, not an optional one. We should be able to talk to LLMs the same way we talk to people — and the system should parse, interpret, and act on that.
This already exists in small ways, but we need more. More tools. More platforms. More experimentation with speech-based interaction.
Imagine coding with your agent — you're typing, but also talking: "Hey, I'm writing a function, generate a test for it and drop it in the test folder." The model sees your context and writes exactly what you need. Or while designing a page: "I want a modern hero section with a signup CTA, spaced out like we did in Project X."
We should also be building LLM-powered systems that support voice-first use cases — AI interviews (like what I'm building with Truefit), AI companions, therapists, and support agents. Voice interaction helps models pick up human nuance: how we speak, how different people pronounce things, how meaning shifts with tone or culture.
More voice input means better data. Better data means better alignment. This is how we move forward.