undefined | Better HN

undefined | Better HN

0 comments

TheEzEzzOP2y ago

Since this is getting a bit of interest, here's one more demo of this https://youtu.be/cvKUa5JpRp4 This demo shows even lower latency, plus the ability to handle very large menus with lots of complicated sub-options (this restaurant has over a billion option combinations to order a coffee). The latency is negative in some places, meaning the system finishes predicting before I finish speaking.

arcticfox2y ago

Holy cow. That's better than the average human drive-through attendant.

jonplackett2y ago

This is cool. But I want to see how it handles you going back one and tweaking it.

armini2y ago

We've built something similar that allows you to tweak/update notes & reminders https://qwerki.com/ (private beta) here's the video demo https://www.youtube.com/shorts/2hpBTxjplIE we've since moved to training our own LLAMA as it's more responsive & we have better reliability.

This is pretty good. Do you think running models locally will be able to achieve performance (getting task done successfully) compared to cloud based ones.i am assuming for context of a drive through scenario it should be ok but more complex systems might need external infromation

TheEzEzzOP2y ago

Definitely depends on the application, agreed. The more open ended the application the more dependent it is on larger LLMs (and other systems) that don't easily fit on edge. At the same time, progress is happening that is increasing the size of LLM that can be ran on edge. I imagine we end up in a hybrid world for many applications, where local models take a first pass (and also handle speech transcription) and only small requests are made to big cloud-based models as needed.

wordpad252y ago

Can you share the source code? What did you do to improve the latency?

TheEzEzzOP2y ago

Lots of work around speculative decoding, optimizing across the ASR->LLM->TTS interfaces, fine-tuning smaller models while maintaining accuracy (lots of investment here), good old fashioned engineering around managing requests to the GPU, etc. We're considering commercializing this so I can't open source just yet, but if we end up not selling it I'll definitely think about opening it up.

Breza2y ago

Neat! I appreciate your approach to preventing hallucinations. I've used something similar in a different context. People make a big deal about hallucinations but I've found that validation is one of the easier aspects of AI architecture.

nelox2y ago

The voice does not seem to be able to pronounce the L in “else”. What’s happening there?

TheEzEzzOP2y ago

Good question. Off the shelf TTS systems tend to enunciate every phoneme more like a radio talk show host rather than a regular person, which I find a bit off putting. I've been playing around with trying to get the voice to be more colloquial/casual. But I haven't gotten it to really sound natural yet.

This is a very slick demo. Nice job!

TheEzEzzOP2y ago

Thanks! It's a lot of fun building with these new models and recent AI approaches.

Wow, the latency on requests feels great!! I’m really curious: is this running entirely with Python?

TheEzEzzOP2y ago

100% Python but with a good deal of multiprocessing, speculative decoding, etc. As we move to production we can probably shave another 100ms off by moving over to a compiled system, but Python is great for rapid iteration.

Manna v0.7

kortex2y ago

Context for the unaware: https://en.m.wikipedia.org/wiki/Manna_(novel)

That's way slick.

Can I ask what your background is, and what things you're used to working with? I don't have the chops to build what you built, but I'd love to get there.

TheEzEzzOP2y ago

My advice is always to jump in and start building! My background is math originally, so I had some of the tools in my tool box, but I'm mostly self-taught in computer science and machine learning. I read textbooks, research papers, code repos, but most importantly I build a lot of stuff. Once I'm excited about an idea I'll figure out how to become an expert to make it a reality. Over the years the skills start to compound, so it also helps that I'm an old man!

simian19832y ago

That demo is pretty slick. What happens when you go totally off book? Like, ask it to recite the numbers of pi? Or if you become abusive? Will it call the cops?

TheEzEzzOP2y ago

It's trained to ignore everything else. That way background conversations are ignored as well (like your kids talking in the back of the car while you order).

edge172y ago

How do you train for this?

yarone2y ago

Nice work, very cool!

j / k navigate · click thread line to collapse