Whisper-asr

Whisper-asr

Docker app from hernandito's Repository

Overview

GPU accelerated Whisper ASR Webservice using the faster-whisper engine, exposing an HTTP API that returns word-level timestamps. Models are stored on the host (ASR_MODEL_PATH) so they are NOT baked into the docker image.

Release Docker Pulls Build Licence

Whisper ASR Box

Whisper ASR Box is a general-purpose speech recognition toolkit. Whisper Models are trained on a large dataset of diverse audio and is also a multitask model that can perform multilingual speech recognition as well as speech translation and language identification.

🎉 Join our Discord Community! Connect with other users, get help, and stay updated on the latest features: https://discord.gg/4Q5YVrePzZ

Features

Current release (v1.10.0) supports following whisper models:

Quick Usage

CPU

docker run -d -p 9000:9000 \
  -e ASR_MODEL=base \
  -e ASR_ENGINE=openai_whisper \
  onerahmet/openai-whisper-asr-webservice:latest

GPU

docker run -d --gpus all -p 9000:9000 \
  -e ASR_MODEL=base \
  -e ASR_ENGINE=openai_whisper \
  onerahmet/openai-whisper-asr-webservice:latest-gpu

Cache

To reduce container startup time by avoiding repeated downloads, you can persist the cache directory:

docker run -d -p 9000:9000 \
  -v $PWD/cache:/root/.cache/ \
  onerahmet/openai-whisper-asr-webservice:latest

Key Features

  • Multiple ASR engines support (OpenAI Whisper, Faster Whisper, WhisperX)
  • Multiple output formats (text, JSON, VTT, SRT, TSV)
  • Word-level timestamps support
  • Voice activity detection (VAD) filtering
  • Speaker diarization (with WhisperX)
  • FFmpeg integration for broad audio/video format support
  • GPU acceleration support
  • Configurable model loading/unloading
  • REST API with Swagger documentation

Environment Variables

Key configuration options:

  • ASR_ENGINE: Engine selection (openai_whisper, faster_whisper, whisperx)
  • ASR_MODEL: Model selection (tiny, base, small, medium, large-v3, etc.)
  • ASR_MODEL_PATH: Custom path to store/load models
  • ASR_DEVICE: Device selection (cuda, cpu)
  • MODEL_IDLE_TIMEOUT: Timeout for model unloading

Documentation

For complete documentation, visit: https://ahmetoner.github.io/whisper-asr-webservice

Development

# Install poetry v2.X
pip3 install poetry

# Install dependencies for cpu
poetry install --extras cpu

# Install dependencies for cuda
poetry install --extras cuda

# Run service
poetry run whisper-asr-webservice --host 0.0.0.0 --port 9000

After starting the service, visit http://localhost:9000 or http://0.0.0.0:9000 in your browser to access the Swagger UI documentation and try out the API endpoints.

Credits

  • This software uses libraries from the FFmpeg project under the LGPLv2.1

Requirements

Requires the unRAID "Nvidia Driver" plugin and an NVIDIA GPU. First start downloads the model (large-v3 ~3 GB) to the Models path below. On a 6GB card keep ASR_QUANTIZATION=int8.

Download Statistics

2,273,233
Total Downloads
122,164
This Month
109,140
Avg / Month

Total Downloads Over Time

Loading chart...

Related apps

Details

Repository
onerahmet/openai-whisper-asr-webservice:latest-gpu
Last Updated2026-08-09
First Seen2023-04-26

Runtime arguments

Web UI
http://[IP]:[PORT:9000]/docs
Network
bridge
Shell
bash
Privileged
false
Extra Params
--gpus all

Template configuration

WebUI / API PortPorttcp

HTTP API + Swagger docs (http://host:9000/docs)

Target
9000
Default
9000
Value
9000
ASR_ENGINEVariable

Engine: openai_whisper, faster_whisper, or whisperx. Use faster_whisper.

Default
faster_whisper
Value
faster_whisper
ASR_MODELVariable

Model: tiny, base, small, medium, large-v3. Larger = more accurate + more VRAM.

Default
large-v3
Value
large-v3
ASR_QUANTIZATIONVariable

Model precision: float32, float16, int8. On GPU this DEFAULTS to float32 (huge). Use int8 so large-v3 fits in ~2.5GB VRAM (needed for 6GB cards like a GTX 1660 Ti); use float16 on bigger cards for max accuracy.

Default
int8
Value
int8
ASR_MODEL_PATHVariable

Where models are stored INSIDE the container. Must match the Models host-path mapping below so models persist on disk instead of bloating the docker image.

Default
/data/whisper
Value
/data/whisper
ASR_DEVICEVariable

cuda (GPU) or cpu.

Default
cuda
Value
cuda
ModelsPathrw

Host folder that persists downloaded models. Keeps the model OUT of the docker image.

Target
/data/whisper
Default
/mnt/cache/appdata/whisper-asr
Value
/mnt/cache/appdata/whisper-asr
NVIDIA_VISIBLE_DEVICESVariable

Which GPU(s) to expose. 'all', or a specific GPU UUID from the Nvidia Driver plugin.

Default
all
Value
all
NVIDIA_DRIVER_CAPABILITIESVariable

GPU driver capabilities to expose.

Default
all
Value
all
MODEL_IDLE_TIMEOUTVariable

Seconds of inactivity before the model is unloaded from VRAM (0 = keep loaded). Set e.g. 300 to free the GPU between jobs at the cost of a reload delay.

Default
0
Value
0