
Surogate Speech: What Its Local Romanian Models Offer
Surogate Speech pairs 116M-parameter recognition with Romanian text-to-speech and three voices. Explore local hardware options and noncommercial terms.
Surogate Speech is releasing small models for the two sides of a voice interface: understanding spoken Romanian and reading Romanian text aloud. The team's September 25 introduction presents Jackrabbit for recognition and Amami for synthesis. Their downloadable model cards make the hardware options and license terms worth a closer look.
- Jackrabbit's Romanian recognition model has 116 million parameters.
- A separate Jackrabbit variant supports streaming input.
- Amami provides three fixed Romanian voices.
- Both model cards specify noncommercial terms under CC-BY-NC-4.0.
What does Surogate Speech recognize?
Jackrabbit's documentation describes an offline FastConformer model with TDT and CTC decoding options. It produces capitalization and punctuation and has CPU-capable execution paths. The package name includes 110M, while the model card gives the fuller parameter count as 116M.
The authors explicitly describe gaps in their evaluation: telephone audio, overlapping speakers, heavy noise and strong regional accents are not covered by the reported tests. Jackrabbit transcribes speech; it does not identify individual speakers. Those boundaries are useful information for anyone designing a local transcription experiment.
How does Amami turn text into speech?
The Amami model card lists a 357M-parameter model with the voices Doina, Tudor and Radu. Packages cover Linux CPU, NVIDIA GPU and Apple Silicon environments. The release uses fixed voices rather than cloning a supplied speaker.
Its documentation also notes that unusual names, acronyms and long codes can be misread. We would evaluate those examples directly in the intended application, listening to complete outputs rather than judging a voice only from a short demonstration. The provider's published benchmark results remain self-reported; we are not presenting an independent performance comparison.
Who can use the downloadable models?
The model cards direct commercial users to Invergent. Downloadable weights therefore should not be mistaken for unrestricted commercial permission. That distinction belongs alongside hardware requirements when assessing new AI tools.
Our reading is that language-specific models can make experimentation more approachable. A developer could test a Romanian reading assistant or transcribe a personal recording without beginning with a large multilingual deployment. The actual experience would still depend on the microphone, runtime and target hardware.
Compared with the different voice-design approach in our Gemini text-to-speech article, this release offers inspectable local components. Its immediate value is a concrete package to evaluate, with explicit language, voice and licensing boundaries.
Sources: Surogate Speech team introduction — September 25, 2026; Jackrabbit Romanian model card and Amami Romanian model card — accessed September 27, 2026. These are the developers' own descriptions and measurements.
More AI Stories

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Meta Muse Gadgets SDK: Build Your Own AI Agent Hardware
Meta's open-source Muse Gadgets SDKs bring its Muse agent to Raspberry Pi and ESP32 builds, and 5,000 free Home Link dongles are going to subscribers.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.
