Anti-Vocale

com.antivocale.app
by Zapstore _@zapstore.dev

Republished from GitHub / F-Droid by the Zapstore main account.

## Full Description (4000 chars max) Too many voice notes and no time to listen? Turn them into text and read them whenever you want. Anti-Vocale transcribes the voice messages of WhatsApp, Telegram, and any other app directly on your phone. It works on recorded calls and audio files too. Everything happens on your device: the audio never leaves your phone. HOW IT WORKS 1. Long-press a voice message and share it with Anti-Vocale 2. Read the transcript that arrives in the notification 3. Copy it or send it back to the chat with one tap PICK WHAT FITS YOU - 99 languages out of the box (Whisper Turbo): the balanced choice for everyone - Fast and light (Parakeet TDT, recommended): 25 European languages, two sizes (862MB high quality, 640MB compact) - Text in real time (Nemotron 3.5): 45 languages, words appear while they are spoken - One compact engine, 59 languages (Qwen3-ASR), including Hindi and Arabic - Missing your language? Community models install in two taps (German, Spanish, Russian, Persian, Ukrainian, Hebrew and more), and you can import your own model file from a link or a folder: every file is integrity-checked on import UNDERSTAND IT, NOT JUST READ IT - AI summaries: long transcripts get a short, readable recap you can copy and share (needs a Gemma model; the prompt is customizable) - Who said what: conversations get speaker labels (SPEAKER 1 / SPEAKER 2) in the transcript and in every export - Any length works: long audio is split and stitched automatically, and if a run fails halfway it keeps the text it already produced - Export as subtitles (SRT, VTT) or timestamped text EVERYDAY COMFORTS - Smart notifications: page through long transcripts without opening the app, one-tap copy or send back - History with search: find any sentence again, retry with a different model - Auto-copy and auto-save to a folder you pick (Drive, Syncthing) - Video files too: the audio track is extracted and transcribed - Faster by design: silence stripping, parallel processing, and automatic use of your phone's AI chip when available - Per-app settings and Tasker automation PRIVACY, FOR REAL No account, no ads, no tracking, no data collection. After you download a model it works fully offline. Your audio stays on your phone, always. OPEN SOURCE Apache 2.0. Source code: https://github.com/RisorseArtificiali/anti-vocale REQUIREMENTS Android 8.0+ | 4GB+ RAM recommended | 326MB to 4.2GB of storage depending on the model

First release: Aug 4, 2026, 12 total releases.

Most recent release: Sep 26, 2026.

Repo

Appears in 0 app stacks.

0 sats / 0 zaps received in the past year.

Sats Received

Underlying data available via MCP: app_zaps, app_releases.

Zap Count

Underlying data available via MCP: app_zaps, app_releases.

Releases

  • Sep 26, 2026 1.13.2
    ## Micro release: one fix A History search that matched nothing used to replace the whole tab, the search field included, leaving no way back to your transcripts short of restarting the app. The search field is now always visible with a clear (X) button, and clearing the history also clears the search. Reported by a daily user on r/droidappshowcase; thank you. **Full changelog**: https://github.com/RisorseArtificiali/anti-vocale/compare/v1.13.1...v1.13.2
  • Sep 22, 2026 1.13.1
    Micro release: low-memory alerts get a direct link to the setting, and the memory guard no longer blocks model loading by default (protection is opt-in). - Memory-failure alerts carry an "Open setting" action that lands on the exact Settings row - Models always try to run; the memory guard is now opt-in, off by default (Settings > Advanced)
  • Sep 21, 2026 1.13.0
    # Anti-Vocale 1.13.0 ## For everyone - Speaker labels in transcripts: two-speaker calls show per-turn speaker annotations in the History detail, with the speaker count detected automatically (GH #83). - Dual-model transcription: a fast streaming first pass renders immediately, then a stronger model refines the transcript in place (GH #43). - Settings search: type a few letters to jump to any option (GH #98). - AI summaries now cover hour-long recordings (map-reduce over the context limit). - Sentence-timed SRT and VTT exports, also for streaming models (GH #92). - Persian interface, the first Persian transcription model, and a text size setting for reading surfaces.
    More…
    - A failed long run keeps and shows whatever transcript it produced, with decoded-of-total context. - Repetition-loop guard: runaway decodes on long runs are detected and cut (GH #107 follow-up work landed). - The Gemma engine moved to LiteRT-LM 0.17.1: the 0.16.1 audio-transcription crash is gone (GH #64). - The History list loads long transcripts much faster: heavy per-row text now loads on expand instead of riding the whole list (TASK-595/599). - The language chip appears only on models that can actually detect the language (Whisper, SenseVoice), and SenseVoice's detected language is stored as a clean code (GH #114). The technical processing line on entries is now opt-in from Settings. - Community catalog: curated per-language profiles (GH #70) and the Orukeet Italian fine-tune. - Every MediaCodec-decodable audio format is accepted up front (GH #18). ## For developers - sherpa-onnx stays at v1.13.8 (srclib pin 11afbd00); the reproducible-build docs and reference workflow we contributed were merged upstream (k2-fsa/sherpa-onnx#3963, tracked in GH #65). - litertlm-android 0.17.1 bumps kotlinx-coroutines-core to 1.11.0 by transitive resolution over our direct 1.7.3; the full suite and the on-device audio E2E ran green on the resolved graph. - Room schema v12 (segments, failure context, processing context, detected language columns). - LlmManager gained a per-generation ceiling that cancels the native stream on timeout, with a wedge breaker (TASK-594); crash-recovery seeding covers every phase-2 arm (TASK-602). - Measured per-model memory footprints now drive the pre-flight check (TASK-575); loop-detector values are persisted for threshold tuning (TASK-582). - Closed in this release: GH #18, #43, #64, #65, #83, #96, #114 and the internal TASK batch 512-616.
  • Sep 16, 2026 1.12.1
    # v1.12.1 Maintenance release: the share-path fixes from Tim Veles's reports plus the model-import and subtitle-choice improvements that landed after v1.12.0. ## For everyone - **Call-recorder shares work again.** Audio shared directly from call-recorder apps (ACR Phone and similar) was rejected as "format not compatible" when the share carried no file name and a generic MIME type. Anti-Vocale now reads the file's first bytes and recognizes the container itself: MP3, M4A/MP4, AAC/ADTS, OGG/Opus, WAV, FLAC, AMR, WebM, 3GP/3G2, MOV (GH #95). - **Imported model folders are auto-detected.** Folder imports no longer ask which family the model is when the folder name or plan files make it clear; the chooser appears only for genuinely ambiguous layouts (GH #93). - **The subtitle-choice timeout is a setting.** Videos with text tracks offer "use subtitles or transcribe audio" and previously always waited 5 minutes before falling back to transcription; the wait (1, 2, 5 or 10 minutes) is now in Settings, Advanced. - **Rejected recordings show their real length.** Entries that fail the duration ceiling displayed 0:00 in History even though the error message quoted the real limit; the actual duration is now recorded on both the whole-file and streaming paths.
    More…
    - **Sharing is faster and safer.** The shared file's copy, the subtitle probe and the preference reads moved off the UI thread (large videos no longer freeze the share transition), and rotating the screen during a share no longer starts a second copy and a duplicate transcription. - **Quieter failure paths.** A share that cannot be enqueued (foreground-service restriction with notifications unavailable) now says so instead of silently doing nothing, and the same honesty applies to the notification action taps. ## For developers - `AudioFormatSniffer` (util, pure Kotlin): magic-byte container detection, 15 JVM tests pinning header bytes; `SharedAudioHandler` treats MimeTypeMap's generic "bin" answer for application/octet-stream as unresolved so the sniffer is reachable, and fills the 16-byte header with a read loop. - MPEG sync-word discrimination now checks version and layer fields: ADTS AAC maps to the supported `aac` extension, reserved-field garbage is rejected; MP4 brand table resolves ordinary video brands (isom/mp42) and 3g2 as video, M4A/M4B/M4P as audio. - `ShareReceiverActivity` runs its flow on the injected `@ApplicationScope` via a new `BackendRegistryEntryPoint.applicationScope()` accessor; recreation is guarded by a pid-stamped saved-state marker (in-process recreation skips, relaunch after process death recovers), the external-model chooser awaits `validRecords()` instead of `runBlocking`, and a chooser that cannot attach to a dead activity drops the share with a toast and an error notification. - `InferenceEnqueue` reports `Failed` instead of pretending to post when notifications are disabled or the fallback channel is blocked; every caller now surfaces the failure (the subtitle tap receiver was the last Log.e-only consumer). - `TranscriptionOrchestrator.failureWritebackSeconds`: one duration writeback rule shared by all three failure catches on both decode paths. - Review follow-ups that did not fit this release are tracked in TASK-524 (items 7-12): the share-URI grant race on the async copy, generic-brand audio-only MP4s badged as video, activity retention through GB-scale copies, and three smaller cleanups. Closes GH #95, closes GH #93.
  • Sep 14, 2026 1.12.0
    # 1.12.0 A big one: a first-launch tour, a new way to feed the app audio, a new language, and the two hardest crash classes of the last months fixed. ## For everyone - **First-launch guided tour.** New installs get a short tour over the real interface: where models live, where results arrive, how to pick a file directly. It plays once; you can replay it from Settings, and a transcription arriving mid-tour dismisses it. - **Pick files directly.** A button on the History tab opens the system file picker. Audio or video, no messaging app needed. Videos with embedded subtitles still offer the subtitle extraction choice. - **Hebrew interface.** The app now speaks 11 languages. The interface fully mirrors for right-to-left, including icons. - **AI summaries grew up.** The summary is copyable, retryable, and you can write your own prompt for it. Entries that cannot have a summary now say why instead of staying silent.
    More…
    - **Settings reorganized.** Options that only apply with certain models or setups now appear only when they can do something. The punctuation pass is new: it uses Gemma to add punctuation to models that do not emit it. - **Fixed: very long history entries could freeze the list.** A degenerate transcription (a model repetition loop) rendered at full height and could make History unusable. Long entries are now bounded, with a note showing how much text is hidden. - **Fixed: history timestamps were not aligned** across conversation groups. - **Fixed: Telegram voice notes showed a wrong duration** (Opus-in-Ogg duration is now read from the granule position). - **Fixed: low-memory and corrupt-model failures now give a clear message** instead of a silent crash, and a corrupt model directory is healed automatically at load time. - **New models in the community catalog**: SenseVoice Small (multilingual, emits uppercase English), Whisper tiny, and curated per-language recommendations. Download dialogs now state the audio-length capability and the disk/RAM fit before you commit to a big download. - **Community icon designs.** Six launcher icon concepts from a contributor are now selectable in Settings, faithfully re-exported as raster layers. ## For developers - sherpa-onnx pinned to 1.13.8 (all four sync points updated; device-verified byte-identical Parakeet output and improved Qwen3). - The six hand-built enqueue sites were consolidated into one `InferenceEnqueue` path that owns the API 31+ background-start restriction; the Tasker trampoline fallback preserves the request instead of dropping it. - Video-with-subtitles probes and the choice notification are shared between the share receiver and the History browse flow (`SubtitleChoice`), keyed on file path so re-offers replace. - Every transcript-rendering surface goes through one shared cap (`CappedTranscriptText`); search highlighting now matches on the original string via `indexOf(ignoreCase)` (the old lowercase-index slicing crashed on Turkish İ). - `isLlm` backend checks route through one predicate in `BuiltInBackendIds`, pinned by a test; a comparison against `DEFAULT_TRANSCRIPTION_BACKEND` had silently inverted for six weeks when the default changed. - The summary and punctuation passes gate on the exact runtime preconditions the orchestrator enforces; the punctuation AUTO mode is dormant (every shipped model punctuates) and is no longer offered in the dropdown. - Transcription engine updated to Android 16 target SDK; foreground service type declared for expedited work; history error reporting has its own notification id inside the reserved-range contract. Assisted-by: Claude <noreply@anthropic.com>
  • Sep 6, 2026 1.11.3
    ## For everyone - **GigaAM v3 (Russian) accuracy fix.** v1.11.2 split long recordings into 3-minute passes, which stopped the native crash (#76) but degraded accuracy well before that: the model was trained on ~25 s utterances and its output degrades on long passes. v1.11.3 processes audio in 30 s passes instead, the same window the other chunked models use. - Measured on device (Realme RMX3853, Android 16), same 3-minute Russian lecture, same model files: v1.11.2 made one 180 s pass and produced garbled text at **60.5% word error rate**; v1.11.3 processes the same audio in **6 chunks of 30 s** at **15.4% WER**, with a complete transcript (methodology and transcripts: #84). - Desktop eval across six Russian lectures: macro WER **65.1%** with 180 s passes vs **10.7%** with 25 s passes; Parakeet is flat across pass lengths, so this is a GigaAM-specific limit, not a pipeline one. - No other model changed. If you do not use GigaAM v3, this update is a no-op. ## For developers - Scope is deliberately minimal: the only change vs v1.11.2 is the bundled-catalog `chunkDurationSeconds` for `gigaam` (180 to 30), its test pin, and documentation. Cut from a branch off v1.11.2; everything else on main (punctuation pass #81, rawTranscript logging, VAD fallback) waits for 1.12.0.
    More…
    - `maxAudioDurationSeconds` was NOT backported: the flag machinery postdates v1.11.2 and its catalog validator rejects unknown flag keys. - Evidence: device A/B transcripts and logcat in the #84 thread; desktop eval in `docs/research/2026-09-05_gigaam-chunk-length-quality.md`; segmentation landscape (nothing beats fixed 30 s within noise) in `docs/research/2026-09-05_gigaam-segmentation-landscape.md`. - Closes the quality regression discussed in #76; crash fix from 1.11.2 unchanged.
  • Sep 4, 2026 1.11.2
  • Sep 1, 2026 1.11.1
    ## For everyone - **Fixed the crash on long, high-quality recordings for phones with less memory.** Voice messages recorded at 48 kHz or above used to be converted to 16 kHz in one big buffer: on phones with a 256 MB per-app heap limit (common on 12 GB flagships, the limit is a build setting, not a RAM measure) a ~8 minute recording crashed with OutOfMemoryError. The conversion now streams chunk by chunk: measured peak memory drops from ~230-270 MB to ~46 MB on an 8-minute 48 kHz file, regardless of recording rate. Both decode paths (batch and progressive) are covered, and the output is bit-exact with the previous converter. - Community catalog: the Canary entries now carry NVIDIA's official Canary Flash names. ## For developers - Release tooling hardened after the 2026-08-31 F-Droid incident: cross-check skill + `scripts/check-fdroid-release.sh` pin the recipe invariants that drifted silently (sherpa srclib pin vs `.sherpa-version`, issue #38); `new-fdroid-version.py` now syncs the srclib pin from `.sherpa-version`; the reference flow is one command (`release-fdroid-references.sh prepare/finalize`) with explicit gates (signed assets must exist before the fork push; NDK preinstall map checked against the recipe's `ndk:` pins). - CI: the F-Droid reproducible job no longer runs on Play-only dispatches (TASK-414); fdroidserver pinned in the buildserver container. - `extract-release-notes.py`: tr-TR and hi-IN headings never matched the version-heading regex (colon vs whitespace, and a missing "में" in the Hindi pattern); single-section locales had masked both since 1.11.0. Both fixed; every locale's Play note now extracts as a single section.
    More…
    - Full release process documentation: `docs/release-runbook.md` (three release-note artifacts, stale-asset rule, per-branch pipeline polling).
  • Aug 30, 2026 1.11.0
    ## For everyone - **Any audio, any model**: long recordings are split and stitched automatically, with chunk sizes adapted to your phone's free memory. The old 6:20 ceiling on Parakeet is gone. - **A queue you can see**: transcriptions appear instantly as queued, can be cancelled, and every entry shows which model ran and for how long. Long-press for retry, copy, delete. - **Community catalog, one-tap imports**: Swiss German Whisper, German Whisper, streaming Spanish, Russian and Arabic fine-tunes, plus the new **Canary 180M Flash tier** (English, German, Spanish, French) for lighter phones. Anyone can publish their own catalog index and the app can switch to it. - **Import your own models** (Transducer, Whisper fine-tunes, CTC, SenseVoice, Canary) from a folder, a Hugging Face URL, or an entry JSON: role matching, SHA-256 verification, architecture selection, each import gets its own share target. - **Languages in their native names**, and integrity checks on every download. - **Fixed**: Qwen3-ASR failed to load on release builds (since 1.9.0); a memory blow-up that could starve low-RAM phones on long voice messages (peak for a 6-minute file dropped from 5+ GB to ~1.5-1.8 GB, confirmed on the reporting Pixel 5); external Whisper models truncating after 30s; the Swiss German fine-tune decoding garbage after one phrase. ## For developers
    More…
    - New external-model family: **Canary** (`EncDecMultiTaskModel`), silence-aligned 10s chunks via `requiresVadAlignedChunking` on the backend interface, 128 mel bands, one catalog entry per language. - `TranscriptionMemoryPolicy`: RAM-derived chunk caps at the single resolution point, both decode paths; floor clamped to the family cap (sub-30s caps no longer throw). - Model tab now renders from the catalog: streaming transducers, catalog-URL architecture (`TASK-401`), language-driven discovery. - Project declares **AI-assisted, human-owned** development: README, CONTRIBUTING, FAQ; commit provenance via `Assisted-by:` trailers. - Measurement notes published: `docs/research/` (canary chunking, FLEURS quality, int8 quantization). - Full changelog: 199 commits, 278 files. Issues closed in this release: #23 #32 #33 #44 #50 #51 #52 #55 #56 #57 #58 #59 #62 #63 #68.
  • Aug 21, 2026 1.10.0-beta.2
  • Aug 18, 2026 1.10.0
    ## Bring your own models Import custom transcription models from a folder on your phone or from a link, in four architectures (Transducer, Whisper fine-tunes, CTC, SenseVoice). Integrity is verified on import, split-file models are supported, and every imported model gets its own entry in the model list and its own share target. Ideal for languages the built-in models don't cover. See the in-app import guide for the supported formats. ## Also in this release - Model tab redesigned: clearer variant selection for Parakeet and Whisper (Turbo, Distil, Medium, Small), honest speed/quality comparison across all eight engines - Neural acceleration (NNAPI) selectable on every device, with automatic CPU recovery if the accelerator misbehaves - Fixed: crash when selecting the Gemma backend (the issue that made 1.9.2 unusable for some F-Droid users) - Fixed: the high-quality quantized Parakeet variant failing to load on some devices
    More…
    - All transcription engines unified under the hood: smaller app, faster updates, new languages easier to add Full changelog: https://github.com/RisorseArtificiali/anti-vocale/releases
  • Aug 4, 2026 1.9.2