Skip to main content
For clinicians

An experimental speech-support beta for supervised evaluation

DysVoxa is a Windows app for phrase-by-phrase local speech recognition, reviewable text options, and user-triggered synthesized output. It is not therapy, a medical device, an emergency system, or a replacement for a person's preferred AAC.

Where DysVoxa may fit

DysVoxa does not diagnose, treat, train speech, or replace SLP intervention. It may be evaluated as an optional communication workflow in which the speaker reviews or edits text before explicitly triggering a synthesized voice.

Clients who may benefit

  • People with dysarthria or other speech differences that conventional ASR handles poorly
  • People able to review, choose, edit, or type the intended phrase before output
  • People who may also benefit from quick phrases and direct text-to-speech
  • People who retain an established fallback communication method

Clinical contexts

  • Supervised usability and communication-support evaluation
  • Review of recognition alternatives and user editing effort
  • Assessment alongside - never instead of - an existing AAC or fallback method
  • Experimental call-routing tests with a remote partner before any practical use

The review-first workflow

1

1. Record a short phrase

The user controls start and stop. The normal target is around 2–15 words rather than continuous listening.

2

2. Recognize on Windows

Standard uses Parakeet and Advanced uses Whisper locally. Either recognizer can be wrong.

3

3. Review candidates

The app may gather recognition alternatives and saved evidence. A suggestion is not a guaranteed correction.

4

4. Choose, edit, then speak

Only the user's explicit Speak action sends the selected text to the Piper synthesized voice.

Current development stage

DysVoxa is an experimental Windows beta. A functioning workflow exists, while significant safety, timing, provenance, packaging, call-routing, and independent evaluation work remains open.

Present in the project

  • Manual phrase recording and local Parakeet or Whisper recognition
  • A chooser, direct editing, direct text-to-speech, and quick phrases
  • Voice-profile enrollment, saved vocabulary, and phrase examples
  • Local Piper output and optional experimental VB-CABLE routing
  • Optional BYOK cloud suggestions that return text for review

Still being strengthened

  • Late-result handling, confirmation, and stale-result protection
  • Candidate source labels, provenance, quality, and ordering
  • CPU response time and recovery from conflicting actions
  • Packaging consistency and offline verification
  • Remote-listener call tests and broader independent evaluation

Evidence and evaluation

The primary quality measure is word error rate, supported by measures that reflect candidate usefulness, user effort, meaning safety, and latency.

  • Was a useful intended phrase available in the chooser?
  • Did a suggestion help or harm the transcript?
  • Did a critical word, negation, medicine, or number change?
  • How much choosing or editing effort was required?
  • Did the result arrive quickly enough on the reference CPU?
  • Reused recordings and one-speaker findings are development evidence, not proof of general effectiveness.
  • The identified 1,321 multi-word TORGO recordings still represent only eight dysarthric speakers.

Important clinical and safety limits

  • DysVoxa is not a medical device, therapy program, emergency system, or validated general solution.
  • Recognition and every correction suggestion can be wrong, including fluent and confident-looking text.
  • The user must review or edit the text and explicitly press Speak; meaning-critical changes are intended to receive extra confirmation.
  • The current evidence cannot support performance claims by cause, severity, accent, age, or speech pattern.
  • DysVoxa must complement rather than remove a user's established fallback communication method.

Privacy and data handling

Local and Off modes keep the normal speech workflow on the Windows computer. Cloud mode is a separate, deliberate BYOK choice:

  • Parakeet and Whisper process microphone audio locally
  • Piper, profiles, settings, and saved repairs stay local in Local or Off mode
  • Optional Cloud mode may send recognized text and relevant saved text context
  • The cloud correction path does not send raw microphone audio
  • Provider retention, privacy, billing, and service terms apply to Cloud mode
  • Saved API keys are protected locally with Windows Data Protection API facilities

Reference hardware

The CPU-only runtime direction is evaluated against a Windows reference computer; this does not guarantee a particular response time or recognition result.

Windows 10 or 11, 64-bit

Reference processor: Intel i5-8250U class with four cores. No dedicated GPU is intended to be required.

16 GB RAM and audio devices

A microphone and speakers or headphones are required. VB-CABLE is optional for experimental call routing.

Review the current demos

The videos document the current Windows workflow. They are demonstrations, not clinical validation or evidence that the app will work for a particular client.

Questions and testing collaboration

Speech-language pathologists, rehabilitation professionals, AAC specialists, and researchers can ask questions or discuss cautious, consent-based evaluation of the beta.

Donors and sponsors

Fund dysarthria speech recognition that stays free to use.

One-time donations, Indiegogo backing, and organizational sponsorship pay for safety work, Windows packaging, and broader evaluation. SIA DysVoxa is a registered company in Latvia, so gifts may not be tax-deductible.