2 independent customer benchmarksReported by LookDeep Health & Verbali

Wake Word Benchmark: DaVoice vs. Commercial Wake Word Engines vs. openWakeWord

In two independent benchmarks — run and reported by LookDeep Health (health tech) and Verbali (assistive communication) on their own production deployments — DaVoice reached 97.65%–99.25% detection accuracy with 0 false-positive events, compared to 72%–78% detection accuracy for leading commercial wake word engines, and 62%–68.65% for openWakeWord (also written “OpenWakeWord” or “Open Wake Word”), the open-source wake word framework. Results reflect each customer’s specific setup; your results may vary.

On-device wake word detection — also called keyword spotting or hotword detection — is typically scored on two axes: how often it misses the real phrase (false rejection rate, or FRR) and how often it fires on nothing (false acceptance rate, or FAR). See how these customer-reported numbers map to FAR/FRR below.

Key takeaways

  • DaVoice: 97.65%–99.25% detection accuracy, 0 false positives, in both independent benchmarks.
  • Commercial wake word engines: 72%–78% detection accuracy; false-positive rates too high for critical use in both tests.
  • openWakeWord: 62%–68.65% detection accuracy, below both customers’ acceptance thresholds.
  • Both benchmarks were independently run and reported by external customers — LookDeep Health and Verbali — not by DaVoice.
Benchmark #1

LookDeep Health (Healthcare)

lookdeep.health ↗

LookDeep Health is a health-tech company building AI-powered patient monitoring for hospitals and care facilities. LookDeep independently evaluated custom wake words across multiple systems for its clinical deployment. In these environments, false positives are unacceptable — an unwanted wake event can trigger recording or alerts during patient care.

  • DaVoice: 0 false-positive events observed over approximately 30 days of production use on LookDeep’s setup. ✅
  • Commercial wake word engines (leading providers LookDeep evaluated): 72%–78% detection accuracy; the best-performing one still logged approximately 2–3 false-positive events per day under the same setup — this did not meet LookDeep’s acceptance threshold for critical use.
  • openWakeWord: 68.65% detection accuracy on this setup, below LookDeep’s acceptance threshold; false-positive rate was not evaluated here.
DaVoice99.25% ✅
Commercial wake word engines(leading providers evaluated)72%–78%
openWakeWord68.65%
0%50%100%

Definition used by LookDeep Health: A false positive is a wake event when no wake phrase was spoken, counted over the monitored period in the stated environment.

“Once we saw the results, it wasn’t a competition — DaVoice won by a knockout. Every other engine we tested threw off multiple false alarms a day; DaVoice had zero. In a hospital setting, that’s not a marginal improvement — it’s the difference between a product we can deploy and one we can’t.”— Tyler Troy, Co-Founder, LookDeep Health
Reported byTyler Troy, Co-Founder @ LookDeep Health• results received Dec 20, 2024
Benchmark #2

Verbali (Assistive Communication / AAC)

verbali.io ↗

Verbali builds MaTalk AI and VerbaliTalk, augmentative and alternative communication (AAC) apps that help non-verbal children and adults communicate. Verbali’s always-listening AAC experience needs to react instantly to a wake phrase without misfiring — false positives disrupt the conversation the app is meant to support.

  • DaVoice: 97.65% detection accuracy with 0 false-positive events observed on Verbali’s setup. ✅
  • Commercial wake word engines (other companies Verbali evaluated): approximately 75% detection accuracy; false-positive rates were too high for production use across every solution tested other than DaVoice.
  • openWakeWord: 62% detection accuracy on this setup, the lowest of the engines Verbali tested.
DaVoice97.65% ✅
Commercial wake word engines(other companies evaluated)~75%
openWakeWord62%
0%50%100%

Note: Detection-accuracy figures are as reported by Verbali. Verbali reported false-positive rates as too high for production use on every engine tested except DaVoice, without a per-engine daily count.

“We knew we’d made the right choice selecting DaVoice, but it really hit home a few months in — watching it perform live in a very noisy area at ATIA Conference 2026 without ever missing a single wake word. It gave our product a genuine wow effect; people were consistently impressed by how well the wake word activation worked.”— Lori Azerrad, Co-Founder & CTO, Verbali

Both Benchmarks, Side by Side

Detection accuracy — higher is better. Figures are as reported by each customer.

Wake word detection accuracy by engine, across both independent benchmarks
ModelLookDeep HealthVerbali
DaVoice99.25% ✅97.65% ✅
Commercial wake word engines72%–78%~75%
openWakeWord68.65%62%

In both independent, unrelated benchmarks, DaVoice was the only engine tested with 0 false-positive events.

How These Numbers Relate to False Acceptance Rate (FAR) and False Rejection Rate (FRR)

Wake word engines are conventionally scored on two industry-standard metrics: False Rejection Rate (FRR) — how often the engine misses a real wake phrase — and False Acceptance Rate (FAR) — how often it triggers on audio that wasn’t the wake phrase, typically expressed as false alarms per hour of listening. Vendor-reported single-number “accuracy” claims are often viewed with skepticism precisely because they can obscure the FAR/FRR trade-off; that’s part of why the results on this page are reported by external customers, not by DaVoice.

  • Detection accuracy on this page ≈ 1 − FRR. DaVoice’s 99.25% (LookDeep Health) and 97.65% (Verbali) detection rates correspond to a false rejection rate of well under 3% in each deployment.
  • False-positive events on this page ≈ FAR. DaVoice’s 0 false-positive events across ~30 days of continuous production listening in both deployments corresponds to a false acceptance rate of effectively zero for that period — versus the best commercial engine tested logging ~2–3 false alarms per day in LookDeep’s test.

These are customer-reported production figures, not a controlled FAR/FRR sweep across noise conditions and accents — see Methodology & Sources below for what was and wasn’t measured.

Methodology & Sources

Benchmark #1: LookDeep Health

Source
Customer-reported results received Dec 20, 2024
Environment
Clinical / hospital deployment, ~30-day monitored period
OS / runtime
Linux (Python)
Model / version
hey_look_deep_model_28_08122024.py2
Detection threshold
0.99

Benchmark #2: Verbali

Environment
MaTalk AI / VerbaliTalk AAC apps, always-listening deployment

Replication: logs and configs available on request — benchmarks@davoice.io. Results reflect each customer’s setup; your results may vary by device, environment, and thresholds. We will clarify or correct this page if new facts emerge.

Frequently Asked Questions

What is the best wake word detection system?

In two independent customer benchmarks, DaVoice achieved 97.65%–99.25% detection accuracy with 0 false-positive events. LookDeep Health reported 99.25% detection with 0 false positives over ~30 days; Verbali reported 97.65% detection with 0 false positives. Results can vary by device, environment, and thresholds.

How do leading commercial wake word engines compare to DaVoice?

Across both benchmarks, leading commercial wake word engines scored 72%–78% (LookDeep Health) and approximately 75% (Verbali) detection accuracy. False-positive rates were too high for critical use in both tests — LookDeep measured ~2–3 events per day from the best commercial engine, and Verbali found every non-DaVoice engine had unacceptably high false positives.

How does openWakeWord (OpenWakeWord) compare to DaVoice?

openWakeWord, the open-source wake word detection framework, scored 68.65% (LookDeep Health) and 62% (Verbali) detection accuracy — both below DaVoice and below each customer’s acceptance threshold.

Are these independent, third-party benchmarks?

Yes — each was run and reported by an external customer on its own production deployment, not by DaVoice. LookDeep Health and Verbali are unrelated companies in different industries that independently reached the same conclusion. Methodology and parameters are published above so the results can be reviewed or replicated.

Who is LookDeep Health?

LookDeep Health is a health-tech company building AI-powered patient monitoring for hospitals and care facilities. LookDeep evaluated DaVoice against openWakeWord and several leading commercial wake word engines before selecting DaVoice for its clinical deployment.

Who is Verbali?

Verbali builds MaTalk AI and VerbaliTalk, augmentative and alternative communication (AAC) apps that help non-verbal children and adults communicate. Verbali evaluated DaVoice against openWakeWord and leading commercial wake word engines before selecting DaVoice for its always-listening AAC experience.

Is wake word detection the same as keyword spotting or hotword detection?

Yes. “Wake word detection,” “keyword spotting,” and “hotword detection” describe the same on-device technology: continuously listening for a specific word or phrase and activating only when it is heard. DaVoice, openWakeWord, and the commercial engines referenced on this page are all keyword-spotting / hotword-detection systems.

What do false acceptance rate (FAR) and false rejection rate (FRR) mean here?

FRR measures how often an engine misses the real wake phrase; FAR measures how often it falsely triggers on other audio. Detection accuracy on this page corresponds to 1 − FRR, and reported false-positive events correspond to FAR. See How These Numbers Relate to FAR/FRR above for the full breakdown.