一款无窗口后台运行的语音输入工具,通过 鼠标中键 + 右键 长按触发,实时将语音转换为文字并模拟键盘输入到当前光标位置。
- 组合键触发:鼠标中键 + 右键同时按住超过 300ms(可配置)激活语音识别
- 实时转录:说话时文字实时出现,延迟低
- 松开即停:释放任意一个按键立即停止识别并输出结果
- 无感运行:无主窗口,仅系统托盘图标运行
- 增量输出:每次只输出新增文字,避免重复输入
- 可配置:通过
config.toml修改热键、音频设备、模型路径等
工具需要 whisper 模型才能运行,推荐下载:
| 模型 | 大小 | 精度 |
|---|---|---|
| ggml-tiny.bin | ~75 MB | 较低 |
| ggml-base.bin | ~150 MB | 中等 |
| ggml-small.bin | ~500 MB | 较高 |
下载地址:GitHub - ggml-org/whisper.cpp
注:模型文件较大,请确保网络稳定。如果下载困难,可以从 HuggingFace Mirror 下载。
将模型文件放入 models/ 目录。
编辑 config.toml:
[hotkey]
press_threshold_ms = 300 # 中键+右键长按多少毫秒激活
[audio]
input_device = "default" # 麦克风设备,"default" 使用系统默认
sample_rate = 16000
channels = 1
buffer_ms = 20 # 音频缓冲区(毫秒)
[recognition]
engine = "whisper_stream"
whisper_exe_path = "./whisper.cpp/Release/whisper-cli.exe"
model_path = "./models/ggml-tiny.bin" # 修改为你的模型路径
language = "zh" # 语言代码,如 "zh"、"en"
[output]
simulate_typing = true
typing_delay_ms = 5 # 每个字符间隔(毫秒)
[tray]
enabled = true
icon_path = "./icon.ico" # 可选:自定义托盘图标双击 voiceinput.exe 即可。运行后在系统托盘可见,右键可退出。
中键(滚轮)+ 右键同时按住 300ms 以上触发语音识别,松开任意一个按键立即停止。
voiceinput/
├── voiceinput.exe # 主程序
├── config.toml # 配置文件
├── icon.ico # 托盘图标(可选)
├── whisper.cpp/
│ └── Release/
│ ├── whisper-cli.exe # whisper 命令行工具
│ └── *.dll # 依赖的 DLL
└── models/
└── ggml-tiny.bin # whisper 模型文件(需自行下载)
- 音频捕获:使用
cpal库,16kHz 单声道 16-bit PCM - 鼠标监听:使用
rdev库,全局鼠标事件捕获 - 键盘模拟:使用
enigo库,模拟键盘文本输入 - 语音识别:通过子进程调用
whisper-cli.exe(whisper.cpp) - 系统托盘:使用
tao+tray-icon库
Q: 提示 "未找到音频输入设备"
A: 检查 config.toml 中 input_device 是否正确,或设为 "default"
Q: 提示 "识别引擎初始化失败"
A: 确保 whisper-cli.exe 路径和 model_path 配置正确,模型文件存在
Q: 文字输出太慢
A: 使用更小的 whisper 模型(如 tiny),或在配置中减少 buffer_ms
Q: 游戏里无法输入文字 A: 部分游戏有保护机制,可尝试以管理员权限运行
MIT License