Typical ASR
is trained mainly on typical speech
Dysarthric speech may produce an incomplete, unrelated, or confidently wrong transcript.
DysVoxa is being developed for people whose speech is difficult for conventional recognition systems to understand. It listens to one short phrase, may offer alternative interpretations, and speaks only the text the user chooses or edits.
Current status: DysVoxa is not yet proven to work reliably across dysarthric speakers. Recognition and suggestions can be wrong. Keep an established fallback method available.
Listening · phrase 1 of 1
Original recognition
“Can you hear me on the call”
Possible interpretation
Choose, edit, or keep the original
Ready after you press Speak
1. Record
One short phrase
2. Recognize
Local speech models
3. Review
Choose or edit text
4. Speak
Explicit Piper output
Being explored for
Dysarthria can affect the strength, speed, range, or coordination of speech movements, and its effects differ widely. Conventional recognition may miss the speaker's meaning even when its transcript looks confident.
Typical ASR
is trained mainly on typical speech
Dysarthric speech may produce an incomplete, unrelated, or confidently wrong transcript.
Fluent errors
can look reasonable while changing the meaning
A normal confidence score may not reveal that the recognizer misunderstood the speaker.
More than one
interpretation may deserve review
DysVoxa can gather recognition alternatives and saved personal evidence instead of trusting the first result.
No guarantee
some phrases cannot be recovered by software
The user can edit or type the text, choose a quick phrase, cancel, or try the phrase again.
DysVoxa does not assume that every phrase can be recovered. Its goal is to reduce the effort needed to reach the intended text without silently changing the speaker's meaning. The speaker remains in control.
Local speech recognition produces a starting point. DysVoxa may gather other interpretations, but the user reviews the text and explicitly decides what - if anything - should be spoken.
Start recording, speak a short phrase - normally around 2–15 words - and stop when the phrase is complete. DysVoxa is not an always-on listener.
Standard uses Parakeet; Advanced uses Whisper. Both run on the Windows PC, and neither recognizer is guaranteed to be better for every speaker.
The app may combine the original result with recognition alternatives, saved vocabulary, prior corrections, phrase examples, and optional suggestions.
Choose an option, keep the original, edit or type different text, use a quick phrase, or cancel and try again. Suggestions are not guarantees.
Only an explicit Speak action starts Piper text-to-speech. The result is a separate synthesized voice, not a clone of the user’s recorded voice.
DysVoxa is designed around review rather than automatic certainty. Suggestions are possibilities, not guaranteed corrections, and no output begins until the user explicitly presses Speak.
Recognition and correction suggestions do not speak automatically. The user chooses or edits the text and explicitly requests speech.
A changed negation, medicine, quantity, name, place, date, or time can alter a message. Meaning-critical changes are intended to receive extra confirmation.
Parakeet, Whisper, Piper, settings, and voice profiles run or stay on the Windows computer in the intended Local workflow.
Saved names, phrases, vocabulary, corrections, and enrolled recordings can inform later stages. They do not retrain the fixed recognition models.
The current Local corrector uses deterministic rules and saved evidence. It has not shown a general phrase-level accuracy improvement.
Cloud mode requires a deliberate selection and the user’s own API key. It can send recognized text and relevant text context, but not microphone audio.
VB-CABLE can provide a virtual audio path, but remote audibility in Zoom and Teams has not yet passed formal end-to-end testing.
DysVoxa is not an emergency system, medical device, or replacement for a person’s preferred AAC. It remains testing software.
The current product is a review-first Windows beta, not an automatic or validated communication solution. Its boundaries matter as much as its features.
| Area | Current scopeDysVoxa is | Not claimedDysVoxa is not |
|---|---|---|
| Interaction | A phrase-by-phrase workflow, normally around 2–15 words at a time | A continuous ambient listener or an automatic call bot |
| Decision | A chooser where the user can review, select, edit, type, or cancel | A system that silently decides what the speaker meant or speaks automatically |
| Output | Piper text-to-speech using a separate synthesized voice | Voice cloning or playback in the user’s recorded voice |
| Platform | An experimental Windows 10 and 11, 64-bit beta | A current Linux product release; Linux is deliberately deferred |
| Local and cloud | Local speech recognition with optional BYOK cloud text suggestions | A service that uploads microphone audio for recognition by default |
| Calls | Able to route synthesized audio through an experimental VB-CABLE path | Not proven end to end for remote listeners in Zoom or Microsoft Teams |
| Role | A communication-support project intended to complement existing fallbacks | A medical device, emergency system, therapy, or replacement AAC method |
| Evidence | An active, experimental beta with limited development data | A validated general solution for every dysarthric speaker |
Even when a transcript looks fluent
Parakeet and Whisper are different recognizers, and either can return an incomplete, unrelated, or confidently wrong result. A confidence score cannot guarantee that the intended meaning was recovered.
Another option is not a guaranteed correction
Saved vocabulary, prior corrections, acoustic examples, sound-based alternatives, and local or cloud suggestions provide evidence for review. The user still decides what to keep, change, or discard.
Fallbacks remain essential
When the audio and saved evidence do not contain enough clues, software may not recover the phrase. The user can edit or type, use a quick phrase, try again, or switch to an established communication method.
DysVoxa has a functioning project workflow, but important safety, timing, candidate-handling, packaging, call-routing, and validation work remains open. It should be treated as testing software.
Current project
Open work
Not current scope
The maintainers intend their project release to remain free, and no paid product tier is planned. The application source is MIT-licensed, while models, voices, datasets, and VB-CABLE keep their own licenses and terms.
This is the current licensing and product direction, not a claim that the experimental beta is ready for dependable use.
Join testing when availableThe MIT application license does not replace the licenses or service terms that apply to bundled and optional components.
Review privacy informationModels, voices, datasets, and VB-CABLE must be attributed and reviewed separately from the MIT-licensed application source.
Funding supports safety architecture, installed-app testing, packaging, remote call validation, broader evaluation, and the work needed to understand where this experimental beta helps or fails.
Campaign progress
Core goal · stretch ceiling
Raised so far
€0
Goal
€28,575
ceiling €33,719
Safety work, quality measurement, hardware, packaging, and evaluation on the Windows reference computer.
The stretch range is an extra €5,000-€5,500 for additional testing, packaging work, and evaluation.
How every euro is allocated across the campaign.
Strengthening candidate handling, confirmation, CPU scheduling, and speech output.
Keeps the project running - hosting, tooling, and ongoing iteration.
Campaign and payment processing fees on the funding platform.
A reserve for hardware, testing, and unforeseen costs.
After diagnosis, everyday communication became an obstacle: repeating myself on calls, losing momentum in conversations, watching routine interactions become exhausting. With significantly limited mobility, reliable remote communication became essential - not optional.
The approach is pragmatic: measure recognition and candidate quality, protect the speaker's intended meaning, and report limits honestly. DysVoxa remains an experimental beta while that work and broader independent testing continue.
- Igor, Founder, SIA DysVoxa
The project measures more than whether a transcript looks fluent. It examines recognition errors, candidate usefulness, meaning-critical changes, user effort, and whether results arrive quickly enough on the reference Windows computer.
Existing recordings have been reused during development, so positive results are development or one-speaker personalization findings - not proof of general effectiveness. The identified TORGO subset adds 1,321 multi-word recordings, but still represents only eight dysarthric speakers.
The primary measure counts words inserted, removed, or changed compared with the spoken reference.
Evaluation asks whether a useful phrase appeared and how much choosing or editing the user needed.
Tests track harmful suggestions, changed critical words or numbers, and response time on the reference CPU.
Direct answers about recognition limits, user control, local and cloud modes, synthesized speech, call routing, and the beta's current status.
DysVoxa is an experimental Windows beta. Register your interest in testing the phrase-by-phrase workflow, and keep an established communication fallback available at all times.
Reference Windows computer
Intel i5-8250U class · 4 cores · 16 GB RAM
Windows 10 or 11, 64-bit, with a microphone and speakers or headphones. A dedicated GPU is not intended to be required.