Talking to J.A.R.V.I.S. Is Harder Than It Looks

My delusions about J.A.R.V.I.S., part two. This time, it’s about voice input and output.

Watching Tony Stark and J.A.R.V.I.S. talk in the movies, I can’t help thinking that with today’s LLMs and a well-crafted prompt, that level of conversation might actually be possible. Polite at all times, yet occasionally sarcastic with a bit of wit, reading Tony’s intentions and getting things done decisively on its own. In text, that is. Turning that into voice is a slightly different story.

Technically, both directions are possible. We have TTS (Text-to-Speech), and we have dictation. It seems like you could just wire them together, but once you actually try to connect them, you run into a few roadblocks.

First, the part where J.A.R.V.I.S. talks back. TTS itself isn’t a big roadblock. There are a few tricky issues, of course. For example, the pronunciation of all sorts of proper nouns that probably don’t exist in any dictionary. People make that mistake all the time too, but with a person, you tell them how to say it once and they fix it right away. With TTS, fixing it may mean going all the way back to the model training stage, so the feedback loop might not be that fast. The bigger problem is that today’s LLM answers are pretty verbose. You could tell it in the prompt to keep answers short, but finding the sweet spot that makes answers short enough without distorting the information won’t be easy.

Dictation is the bigger problem. There are technical hurdles and non-technical hurdles here. Starting with the technical ones, there are a few, but the first one that comes to mind is multilingual support. Mixing words from multiple languages in a single sentence isn’t rare anymore. Then there are cases like ‘네’ (ne, Korean for “yes”) and ‘Nay’, which sound similar but mean completely different things. In my case, as a developer who mainly speaks Korean, the instructions I give AI often mix Korean and English. For example: “statusline.sh 의 make_bar 폭을 8칸으로 줄여줘.” (“Shrink the make_bar width in statusline.sh to 8 cells.”) The trickier part is the non-technical hurdle. The point is, we’re not Robert Downey Jr., giving J.A.R.V.I.S. instructions from neatly memorized lines. Rewrite the example above more realistically, and it looks like this: “Uh… so… statusline.sh … the make_bar part’s .. um.. length? No.. width..right! Make the width smaller than it is now… hmm.. how much would be good? 6 cells? 8 cells? 8 sounds good. Shrink it to 8 cells.” On top of that, ordinary people like me don’t even pronounce things accurately. I think I said it right, but it’s pretty common for a speech recognition model to hear me as saying something totally different. Like when “Shrink it to 8 cells” gets transcribed as “Shrink it to ate sells.” Even so, I’ve personally been trying out dictation in various ways in everyday life. I’ll write up that trial and error in another post some other time.

To sum up, the smooth conversations between Tony Stark and J.A.R.V.I.S. are only possible because they’re a rehearsed performance. Reality is a lot messier. It’s not technically impossible, of course, but it’s hard to make it as clean as in the movies. Because we’re not actors, just regular folks …

Leave a comment