DEV Community

Cover image for Escape the Keyboard: LocalWhisper Pro Helps You Actually "Touch Grass"
Harsh Bhadu
Harsh Bhadu

Posted on AI-assisted

Escape the Keyboard: LocalWhisper Pro Helps You Actually "Touch Grass"

Hacktoberfest: Maintainer Spotlight

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

LocalWhisper Pro is an open-source, offline-first voice interface designed to minimize screen time. Instead of spending hours hunched over a keyboard typing emails, documentation, or field notes at 40 WPM, users dictate naturally at 150+ WPM. Local AI instantly cleans, restructures, and auto-types the text directly at the cursor—enabling users to wrap up screen work fast and get outdoors.

How it gets people off the screen and into the world:

  • Makes the screen the shortest interaction: Speech is up to 4x faster than typing. By handling filler-word scrubbing, formatting, and summarization automatically, users spend a fraction of the time staring at word processors and code editors.
  • 100% Offline Trail & Field Ready: Because both transcription and refinement run entirely on-device, you can take your laptop outside—onto the porch, into a garden, or down a hiking trail without cell signal—to dictate field logs, nature observations, and ideas without internet dependencies.
  • No Clipboard or Cloud Distractions: It uses low-level hardware stroke injection (SendInput / pynput) rather than pasting through clipboard history, keeping workflows seamless and distraction-free.

Demo

  • Direct Download: Download the standalone Windows executable (LocalWhisper_Pro_v3.1_Windows.zip) directly from our GitHub Releases page—no Python runtime required.

Code

Repository: [https://github.com/Harsh-Dev07/Local-WhisperFlow]
License: Open Source (MIT)


How I Built It

LocalWhisper Pro is architected entirely around open-source AI models and local runtime frameworks:

  1. Local Speech Recognition (faster-whisper / CTranslate2):
    Audio input is captured locally using sounddevice and processed through faster-whisper, a reimplementation of OpenAI's Whisper model running on the CTranslate2 inference engine. This delivers up to 4x faster-than-real-time transcription on consumer CPUs/GPUs without external network calls.

  2. Local Intelligence & Text Refinement (Ollama + llama3.1:8b):
    Raw transcription often suffers from stutters, repetitions, and unstructured ramblings. LocalWhisper integrates directly with local Ollama instances. Using open-weight models like llama3.1:8b, it cleans transcripts across multiple modes:

    • Smart Dictation: Removes filler words while maintaining personal voice.
    • Field & Bullet Summary: Distills stream-of-consciousness rambling into structured, actionable checklists.
    • Code & Tech: Synthesizes voice logs into camelCase, snake_case, and terminal commands.
    • Hinglish/Hindi Processing: Accurately interprets code-switched multi-language dictation.
  3. System Overlay & Direct Keystroke Emulation:
    The interface is a lightweight, non-stealing glassmorphic HUD built in Python (Tkinter). Refined text bypasses the operating system's clipboard (Win+V) and is typed directly into active windows via hardware-level event hooks (pynput), ensuring zero data retention outside the user's encrypted local history.


Why Does Open Innovation Matter?

Open innovation was not an afterthought for this build—it was a prerequisite:

  • True Offline Portability: Closed-source voice-to-text platforms (like Google Speech-to-Text or OpenAI Whisper API) fail the moment you walk into a park, mountain trail, or remote campsite without Wi-Fi or cellular service. Open-weight models (faster-whisper + Llama 3.1) run autonomously on consumer laptops off the grid.
  • Complete Sensory Privacy: Voice notes recorded outdoors or in private moments often include personal thoughts, journal entries, and sensitive project ideas. Open-source local models guarantee that no spoken audio or transcribed text ever hits corporate telemetry servers or third-party training pipelines.
  • Zero Ongoing Cost: Proprietary APIs charge per audio minute and per token, which discourages long, ambient voice captures. Open-weight inference costs nothing to run, whether you record a two-second note or an hour-long outdoor brainstorm.
  • Model Modularity: Users have the freedom to swap out the refiner engine for smaller models (e.g., phi3, gemma2:2b) on low-power devices, or larger parameter models on dedicated hardware.

Prize Categories

  • Hacktoberfest Open-Source AI Challenge: Week 1 (Touch Grass)

Top comments (0)