I'm not so sure, I think it could go the other way. The vast majority of support cases should be handled in an automated way. I had an issue with Vercel recently where I argued that a bill was incorrect, and the agent produced and offered a refund by itself - that was interesting.
You obviously need humans but they can be freed up to deal with the more complicated cases.
I've been waiting for this! I'm learning Spanish so I built an app to teach me Spanish, but hyperfocused on scenarios in my life, for example "watching a Barça match in a Barcelona bar". It does FSRS flashcard training, and live conversation practice.
I think education is a very underexplored area for these live conversation models. Yes you can just use ChatGPT Live but that's freeform and unstructured, doesn't have a curriculum or can present supporting visuals, etc. On a grand scale if you can give children their own personal individual tutor rather than relying on group teaching alone, there could be a huge jump in successful education outcomes.
Oh yeah, I built so much stuff to learn German, for example [1] to give me random German texts, force me to read it, and answer it, I created [2] to automatically make flashcards for me and then use with with a flashcards app I regularly use and [3] to help me memorise German cases and word-genders. I love it!
I did think of implementing this conversationally, but tbh it has always been too expensive thus far, I gotta retry with GPT-live-1, I tried it with elevenlabs before but it wasn't live enough and the models were not intelligent enough.
Yeah my current approach till now has been ElevenLabs Scribe v2 transcription, then feed that to Gemini Live. The latency isn't too bad, but it's definitely there.
When you use ChatGPT Live it's instant which is great, although the realtime transcription still kinda sucks, especially if you're a newcomer to the language so you're making mistakes. I'll constantly get responses to something it thinks I said but I didn't say, which is a real hard blocker for a language learning app.
I think what I'll land on is Scribe v2 (the full thing, not realtime) transcribing turns - it is exceptionally accurate for this - and then just feeding that text direct to GPT Live.
Oh yeah definitely, That being said what helped me more than anything is the flashcards app and speaking German with my roommate regularly.
Nonetheless, in complete honestly I do also have a German tutor who I see once a week for 50 minutes, I am very reliable on completing my work though, the "progressbars" in my flashcards app do keep me motivated.
Learning a language is really hard and takes years, but mentally I am convinced, that if the progressbars in the flashcard app I use reach 100% and also in my German cases app, that I will get closer to speaking perfect German, this keeps me motivated.
Yeah, it's an exciting use case. Although all these models, even seemingly GPT-Live-1 doesn't actually hear your pronunciation, it seems they all get passed transcripts, so for learning to speak another language, they're still not there seemingly.
But, it's close! You can control their pronunciation, make them speak slower/faster, and obviously great at anything text, so many use cases work great for language learning with LLMs. Just wish they solved this last mile thing too!
Some of the models claim to be audio to audio, like one of the Gemini models. But I've tested and it does seem that you're right, it's not getting all the nuance at all
Yeah, sadly "audio to audio" seems to mean "we transcript it automatically for you internally which gets passed to the model", otherwise we'd be seeing models that are able to hear nuance in the input voice and pronunciation, which AFAIK, no model does yet.
the latest Google Translate features based on the all voice real-time 3.5 Live model is as good as I have tried in Live Modes it's not perfect but I can put it down on a table of four or five people conversing and get a reasonable amount of it translated into my earpiece.
Seems like a very popular use for AI, I work on a version for Korean. (A very diffrent feature set, more exercise generation, maybe someday I'll get to live conversation, which would be a great thing.)
However, for voicing sentences I use murf.ai, which seemed very nice for korean.
I tried to talk to the openai version in (my bad) korean, and it responded in japanese :D
Group teaching for language learning is essential, even more for kids. It’s really important to be in a context where you have to interact with other humans. There is a reason anyone serious and with the means will pay good money to go to language courses IRL to progress, instead of relying on video calls. And the last thing kids need is even less human contact during their education
Not entirely joking but group teaching: fire up multiple AIs. Hear live talking in a foreign language, chip in when you want, they adapt to your level.
Lack of human contact and alienation are real problems at all ages, but that does not mean current chatbots are not excellent teachers, especially in language learning when you do not even need a curriculum as such, just talk/read/write as much as you can, it's all text.
It’s not really about the number of people, it’s about having a diversity of individuals, interests, topics of discussion, people making different type of mistakes, different accents and pace, etc
While not my technical first, it was the first of real consequence for me as well.
And then I got so frustrated that I went to write my own forum software! And showed it to my friend, who promptly showed me how badly I understood security by posting as me.
reply