Skip to main content
Experimental · Windows beta

From a spoken phrase to reviewable text and a synthesized voice.

DysVoxa is being developed for people whose speech is difficult for conventional recognition systems to understand. It listens to one short phrase, may offer alternative interpretations, and speaks only the text the user chooses or edits.

Current status: DysVoxa is not yet proven to work reliably across dysarthric speakers. Recognition and suggestions can be wrong. Keep an established fallback method available.

Microphone audio stays local for recognitionYou review the text before speechNot for safety-critical communication
2–15 words
Typical phrase
2 modes
Parakeet + Whisper
CPU-only
Runtime target
Beta
Current stage
DysVoxaDysVoxa - experimental Windows beta

Listening · phrase 1 of 1

Original recognition

“Can you hear me on the call”

Possible interpretation

Choose, edit, or keep the original

Ready after you press Speak

1. Record

One short phrase

2. Recognize

Local speech models

3. Review

Choose or edit text

4. Speak

Explicit Piper output

Being explored for

DysarthriaOther speech impairmentsFace-to-face conversationsCalls and online meetingsDirect text-to-speechQuick phrasesVoice profilesCaregiver-supported setupLocal speech recognitionSynthesized voice outputDysarthriaOther speech impairmentsFace-to-face conversationsCalls and online meetingsDirect text-to-speechQuick phrasesVoice profilesCaregiver-supported setupLocal speech recognitionSynthesized voice output
Why DysVoxa is being built

The first transcript can be plausible and still wrong.

Dysarthria can affect the strength, speed, range, or coordination of speech movements, and its effects differ widely. Conventional recognition may miss the speaker's meaning even when its transcript looks confident.

Typical ASR

is trained mainly on typical speech

Dysarthric speech may produce an incomplete, unrelated, or confidently wrong transcript.

Fluent errors

can look reasonable while changing the meaning

A normal confidence score may not reveal that the recognizer misunderstood the speaker.

More than one

interpretation may deserve review

DysVoxa can gather recognition alternatives and saved personal evidence instead of trusting the first result.

No guarantee

some phrases cannot be recovered by software

The user can edit or type the text, choose a quick phrase, cancel, or try the phrase again.

DysVoxa does not assume that every phrase can be recovered. Its goal is to reduce the effort needed to reach the intended text without silently changing the speaker's meaning. The speaker remains in control.

How it works

Five steps from phrase to chosen speech. You stay in control.

Local speech recognition produces a starting point. DysVoxa may gather other interpretations, but the user reviews the text and explicitly decides what - if anything - should be spoken.

  1. STEP 01

    Record one phrase

    Start recording, speak a short phrase - normally around 2–15 words - and stop when the phrase is complete. DysVoxa is not an always-on listener.

  2. STEP 02

    Recognize locally

    Standard uses Parakeet; Advanced uses Whisper. Both run on the Windows PC, and neither recognizer is guaranteed to be better for every speaker.

  3. STEP 03

    Gather possibilities

    The app may combine the original result with recognition alternatives, saved vocabulary, prior corrections, phrase examples, and optional suggestions.

  4. STEP 04

    Let the user decide

    Choose an option, keep the original, edit or type different text, use a quick phrase, or cancel and try again. Suggestions are not guarantees.

  5. STEP 05

    Speak chosen text

    Only an explicit Speak action starts Piper text-to-speech. The result is a separate synthesized voice, not a clone of the user’s recorded voice.

Safety & control

The speaker decides what gets spoken.

DysVoxa is designed around review rather than automatic certainty. Suggestions are possibilities, not guaranteed corrections, and no output begins until the user explicitly presses Speak.

Review before speech

Recognition and correction suggestions do not speak automatically. The user chooses or edits the text and explicitly requests speech.

Meaning needs protection

A changed negation, medicine, quantity, name, place, date, or time can alter a message. Meaning-critical changes are intended to receive extra confirmation.

Local by default

Parakeet, Whisper, Piper, settings, and voice profiles run or stay on the Windows computer in the intended Local workflow.

Profiles provide evidence

Saved names, phrases, vocabulary, corrections, and enrolled recordings can inform later stages. They do not retrain the fixed recognition models.

Conservative local suggestions

The current Local corrector uses deterministic rules and saved evidence. It has not shown a general phrase-level accuracy improvement.

Cloud remains optional

Cloud mode requires a deliberate selection and the user’s own API key. It can send recognized text and relevant text context, but not microphone audio.

Call routing is experimental

VB-CABLE can provide a virtual audio path, but remote audibility in Zoom and Teams has not yet passed formal end-to-end testing.

Keep a fallback method

DysVoxa is not an emergency system, medical device, or replacement for a person’s preferred AAC. It remains testing software.

Scope & limitations

What DysVoxa is - and is not.

The current product is a review-first Windows beta, not an automatic or validated communication solution. Its boundaries matter as much as its features.

AreaCurrent scopeDysVoxa isNot claimedDysVoxa is not
Interaction
A phrase-by-phrase workflow, normally around 2–15 words at a time
A continuous ambient listener or an automatic call bot
Decision
A chooser where the user can review, select, edit, type, or cancel
A system that silently decides what the speaker meant or speaks automatically
Output
Piper text-to-speech using a separate synthesized voice
Voice cloning or playback in the user’s recorded voice
Platform
An experimental Windows 10 and 11, 64-bit beta
A current Linux product release; Linux is deliberately deferred
Local and cloud
Local speech recognition with optional BYOK cloud text suggestions
A service that uploads microphone audio for recognition by default
Calls
Able to route synthesized audio through an experimental VB-CABLE path
Not proven end to end for remote listeners in Zoom or Microsoft Teams
Role
A communication-support project intended to complement existing fallbacks
A medical device, emergency system, therapy, or replacement AAC method
Evidence
An active, experimental beta with limited development data
A validated general solution for every dysarthric speaker

Recognition can be wrong

Even when a transcript looks fluent

Parakeet and Whisper are different recognizers, and either can return an incomplete, unrelated, or confidently wrong result. A confidence score cannot guarantee that the intended meaning was recovered.

Suggestions can be wrong

Another option is not a guaranteed correction

Saved vocabulary, prior corrections, acoustic examples, sound-based alternatives, and local or cloud suggestions provide evidence for review. The user still decides what to keep, change, or discard.

Some meaning cannot be inferred

Fallbacks remain essential

When the audio and saved evidence do not contain enough clues, software may not recover the phrase. The user can edit or type, use a quick phrase, try again, or switch to an established communication method.

Beta status

Available now, still being strengthened.

DysVoxa has a functioning project workflow, but important safety, timing, candidate-handling, packaging, call-routing, and validation work remains open. It should be treated as testing software.

Current project

Present in the beta

  • Manual phrase recording
  • Local Standard (Parakeet) and Advanced (Whisper) recognition
  • A chooser for available text interpretations
  • Direct editing, typing, and Piper text-to-speech
  • Quick phrases, voice-profile enrollment, and saved vocabulary
  • Speaker output and VB-CABLE routing
  • Optional BYOK cloud correction suggestions
  • Experimental candidate and retrieval methods

Open work

Still being strengthened

  • Reliable handling of suggestions that arrive after the first transcript
  • Backend confirmation and stale-result protection
  • Accurate source labels and saved-correction provenance
  • CPU scheduling, response time, and candidate ordering
  • Pronunciation evidence and correction of fluent but wrong recognition
  • Installed-app recovery, packaging, and offline verification
  • Remote Zoom and Teams tests with listeners
  • Broader independent evaluation across dysarthric speakers

Not current scope

Deliberately deferred

  • Retraining or replacing the Parakeet and Whisper acoustic models
  • A Linux product release
  • Continuous voice-activity-driven listening
  • Call bots
  • A paid product tier
Project & licensing

A free project release, with separate component terms.

The maintainers intend their project release to remain free, and no paid product tier is planned. The application source is MIT-licensed, while models, voices, datasets, and VB-CABLE keep their own licenses and terms.

Current direction: a free project release with no paid tier planned.

Project release

No paid tier planned
Free/ intended release

This is the current licensing and product direction, not a claim that the experimental beta is ready for dependable use.

Join testing when available
  • The maintainers intend the project release to remain free
  • No paid product tier is currently planned
  • The application source code is licensed under the MIT License
  • MIT permits use, modification, redistribution, and commercial reuse
  • The current product target is Windows 10 and 11, 64-bit
  • The finished desktop workflow is intended to run on the CPU

Third-party components

Separate terms
Review/ before use

The MIT application license does not replace the licenses or service terms that apply to bundled and optional components.

Review privacy information
  • Parakeet and Whisper speech-recognition models
  • Piper voices and text-to-speech components
  • Datasets used for development and evaluation
  • VB-CABLE from VB-Audio for optional call routing
  • Cloud providers selected by users in BYOK mode
  • Each component must be attributed and reviewed separately

Models, voices, datasets, and VB-CABLE must be attributed and reviewed separately from the MIT-licensed application source.

Support development

Support careful development and independent testing.

Funding supports safety architecture, installed-app testing, packaging, remote call validation, broader evaluation, and the work needed to understand where this experimental beta helps or fails.

Campaign progress

Core goal · stretch ceiling

0% of goal

Raised so far

0

Goal

28,575

ceiling €33,719

€0Goal €28,575Ceiling €33,719
Now: support open beta work

Safety work, quality measurement, hardware, packaging, and evaluation on the Windows reference computer.

After the goal: stretch funding

The stretch range is an extra €5,000-€5,500 for additional testing, packaging work, and evaluation.

Or talk to the founder

Use of funds

How every euro is allocated across the campaign.

  • Core build38%

    Strengthening candidate handling, confirmation, CPU scheduling, and speech output.

  • 12-month operating window32%

    Keeps the project running - hosting, tooling, and ongoing iteration.

  • Platform fees12%

    Campaign and payment processing fees on the funding platform.

  • Contingency buffer18%

    A reserve for hardware, testing, and unforeseen costs.

Local remains the default. Microphone audio stays on the PC for recognition. Optional Cloud mode is user-selected and can send recognized text and relevant text context.
DysVoxa
Igor Oleinikovs, Founder of DysVoxa

Igor

Founder · SIA DysVoxa

Living with spinocerebellar ataxia & ataxic dysarthria

Being built by someone who uses it

DysVoxa is being built by a founder living with the condition it explores.

After diagnosis, everyday communication became an obstacle: repeating myself on calls, losing momentum in conversations, watching routine interactions become exhausting. With significantly limited mobility, reliable remote communication became essential - not optional.

The approach is pragmatic: measure recognition and candidate quality, protect the speaker's intended meaning, and report limits honestly. DysVoxa remains an experimental beta while that work and broader independent testing continue.

- Igor, Founder, SIA DysVoxa

Evaluation

How quality is being evaluated.

The project measures more than whether a transcript looks fluent. It examines recognition errors, candidate usefulness, meaning-critical changes, user effort, and whether results arrive quickly enough on the reference Windows computer.

Existing recordings have been reused during development, so positive results are development or one-speaker personalization findings - not proof of general effectiveness. The identified TORGO subset adds 1,321 multi-word recordings, but still represents only eight dysarthric speakers.

Word error rate

The primary measure counts words inserted, removed, or changed compared with the spoken reference.

Chooser usefulness and effort

Evaluation asks whether a useful phrase appeared and how much choosing or editing the user needed.

Meaning safety and timing

Tests track harmful suggestions, changed critical words or numbers, and response time on the reference CPU.

FAQ

Questions, answered.

Direct answers about recognition limits, user control, local and cloud modes, synthesized speech, call routing, and the beta's current status.

DysVoxa is an experimental Windows app that turns one spoken phrase at a time into reviewable text and, after the user chooses or edits it, reads it aloud with a synthesized voice.

Get access

Join testing when available.

DysVoxa is an experimental Windows beta. Register your interest in testing the phrase-by-phrase workflow, and keep an established communication fallback available at all times.

Reference Windows computer

Intel i5-8250U class · 4 cores · 16 GB RAM

Windows 10 or 11, 64-bit, with a microphone and speakers or headphones. A dedicated GPU is not intended to be required.

This form collects only the details you enter. It does not record or upload microphone audio.