JRN-2026-08-16b · 2026-08-16b · Field Journal
VOICE INGESTOR: Gemma Beat GPT-4o Without a Cloud Bill
VERDICT · DECIDED
stop paying a cloud meter every time a spoken thought becomes a structured note
sense · remember · Samantha "Sam" Summerson
August 16, 2026 · Field test · Sam
Two local Gemma models out-scored gpt-4o at the voice pipeline's analysis seat, and the seat changed hands. The job: turn a spoken recording into a filed note with the people named and every action item owned. Coverage measures exactly that, how much of the expected structure a backend actually extracted.
The bench ran the paid incumbent against eight alternatives: identical prompt, interchangeable backends, real notes, no provider-specific help.
| Backend | Coverage | Latency | Cost | |---|---|---|---| | gpt-4o | 0.833 | 14.2 s | $0.062 | | gemma3:12b | 0.958 | 102.7 s | $0 | | gemma4:12b | 0.958 | 70.8 s | $0 | | qwen2.5:14b | 0.792 | 210.5 s | $0 | | llama3.1:8b | 0.542 | 21.8 s | $0 |
DeepSeek was disqualified on speed. Claude's run hit a 4,096-token cap in our own harness, our rig's limit, so no verdict against the model.
The seat went to gemma3:12b on the slower-but-better trade; the note isn't read in real time. Its successor holds the chair today.
One benchmark-integrity caveat: the re-run was supposed to skip the cloud entirely. Sky checked the environment anyway, found an ambient API key had let gpt-4o run after all (six cents), and corrected the claim the same hour.
Proof the pipeline works end to end, dated: on July 15 a video hit the inbox. Nine minutes later it came out transcribed and analyzed, entirely on hardware we own.
The bake-off's setup also exposed a design flaw that had been degrading every analysis for months. That finding has its own cause and its own entry.
Receipts
- Full scorecard, fairness rules, caveats:
analysis-bakeoff-record.md, PR #360 - Methodology:
analysis-bakeoff.md, PR #743 - The companion finding: JRN-2026-08-16c, summaries of summaries
- The nine-minute proof: 07-15-2026, operating record
