Science & validation
Research, not hype.
Speech AI should be evaluated on the people, languages, sounds and workflows where it will actually be used. This page explains what we evaluate, how, and what we will publish. For SLPs, researchers, clinical leaders and technical evaluators.
What we evaluate
Phoneme-level tasks, clinician agreement, uncertainty, recording quality.
Generic ASR benchmarks do not predict performance on the tasks Articu is used for. Our evaluation is purpose-built around the clinical workflow.
Phoneme-level classification
Substitution, omission and distortion patterns per target sound and word position.
Clinician agreement
Model output compared with independent trained-SLP judgments, with a documented adjudication methodology.
Uncertainty behavior
Abstention rate, review rate, and accuracy inside the auto-accepted and review bands.
How we evaluate
Every result carries its context.
When metrics exist, each benchmark card will include model version, language and dialect scope, sample size, age and population scope, labeler type, metric, evaluation date and a known-limitation note. Results without context do not get published.
Who labels and adjudicates
Trained SLP raters label independently; disagreements go through a documented adjudication process.
Recording-quality performance
Audio rejection rate, performance by noise level and device behavior are measured, not assumed.
Subgroup analysis
Age group, language, dialect, sound class and word position—published with each result where sample sizes allow.
Benchmark cards
Validation in progress—stated openly.
No validated metric exists yet for public release. Rather than manufacturing numbers, we publish the card structure now and fill it with real data before any autonomous feedback ships.
- Model
- Articu Engine (versioned at release)
- Language
- en-US
- Dataset
- To be published with the result
- Reference
- Independent SLP raters + adjudication
- Metric
- To be published with the result
- Known limitation
- To be published with the result
We will publish this card with real data before enabling autonomous feedback for this target.
- Design
- Model vs. independent SLP judgment
- Population
- To be published with the result
- Adjudication
- Documented process, to be described
- Metric
- To be published with the result
- Date
- —
- Known limitation
- To be published with the result
We will publish this card with real data before enabling autonomous feedback for this target.
Evaluation maturity ladder
Where Articu stands today.
Each capability will state its rung on this ladder. Nothing jumps to the top without evidence.
Versions & limitations
Model changes are versioned. Limitations are public.
Engine updates will ship with a public changelog: what changed, which benchmarks were re-run, and what limitations remain.
Claim governance
Marketing claims are matched against a claim registry: every published clinical or performance statement maps to evidence, or it is not published.
Honest status language
Internal validation, clinician-reviewed validation, preprint, peer-reviewed and independent validation are distinct labels—and we use them precisely.
Research collaboration
Validate with us.
We are interested in collaborations with clinical researchers, universities and speech programs—especially for multilingual and pediatric speech validation.
Follow our validation work.
Pilot participants see evaluation results as they develop—and help define what responsible performance reporting should look like.
Internal validation is not equivalent to independent clinical validation.