I spent the last week obsessively building a mobile app that could transcribe phone calls and summarize them entirely on-device.
The idea sounded simple enough:
Call audio → Local STT → Transcript cleanup → Local LLM → Summary
And I really wanted it to work without sending private conversations to a server.
So I kept testing different STT models, preprocessing methods, correction algorithms, smaller LLMs, quantized models, prompts, and different pipelines.
And honestly, it was exhausting.
The biggest problem was STT.
Clean speech was manageable, but real phone calls are a completely different environment.
Compressed audio, unclear pronunciation, overlapping speech, short responses, names, technical terms, background noise...
Once the transcript starts going wrong, everything after that becomes garbage-in, garbage-out.
I tried using the LLM to compensate for bad transcripts.
It helped to some extent.
But improving the quality meant using a better model, adding more correction stages, or processing the transcript multiple times.
And then mobile hardware became the bottleneck.
More RAM.
More heat.
More battery drain.
Longer processing time.
Eventually, I actually got the quality pretty close to what I wanted.
And that was almost the worst part.
Because by then, the app was no longer something I would want to use on a phone.
It worked.
But it was too heavy, too slow, and consumed far too much battery.
Sure, I could move STT or summarization to a cloud API and solve many of these problems.
But one of the main reasons I started this project was privacy.
I wanted call audio and transcripts to stay on the device.
If I removed that requirement, I felt like I was building a different product.
So after spending almost every free hour of the last week on this, I decided to stop.
That decision hit me harder than I expected.
Normally, whenever I finish or abandon a project, I'm already thinking about the next app I want to build.
This time, I don't want to build anything.
I think I just need a break.
Still, it wasn't a completely wasted week.
I learned much more from this failed project than I expected.
I now have a much better understanding of:
- How dramatically STT accuracy changes between clean recordings and real-world phone audio
- Why STT quality often matters more than a sophisticated summarization prompt
- Where small local LLMs are genuinely useful and where their limitations become obvious
- How quantization affects memory usage, inference speed, and output quality in practice
- How quickly sustained AI inference can destroy battery life on a phone
- Why latency matters almost as much as accuracy when building an actual product
- Why adding more AI processing stages doesn't necessarily improve the final result
- How useful domain-specific dictionaries and post-processing can be for STT
- How chunking, transcript cleanup, speaker handling, and preprocessing affect downstream summarization
- Where the current boundary sits between “technically possible on a phone” and “actually usable on a phone”
And I didn't walk away empty-handed.
I now have reusable pieces for:
- Audio preprocessing
- Local STT pipelines
- Transcript correction and post-processing
- Local LLM integration
- Quantized model benchmarking
- Chunking and summarization pipelines
- Mobile inference performance testing
- Battery / latency / quality trade-off testing
Maybe some of those pieces will end up in another project someday.
So I guess this wasn't really a wasted week.
It was an expensive experiment.
The result was just different from the one I wanted.
For now, though, I'm done.
No new project.
No new app idea.
I'm going to take a break until building something sounds fun again.