SceneWhisper

SceneWhisper

Docker app from Master-Oogwayxxx's Repository

Overview

AI-powered scene analysis and captioning for local video libraries. Includes a React web UI, FastAPI worker, ffmpeg processing, and optional NVIDIA GPU acceleration.

SceneWhisper

SceneWhisper is a Docker Compose packaged video processing app with a Python FastAPI backend worker and a browser-based frontend service. This repository is organized for Unraid and Docker deployment, including Community Applications template files and a GHCR release workflow.

Project Layout

SceneWhisper/
├── app/
│   ├── backend/              # Python 3.11 FastAPI API and worker
│   │   ├── src/
│   │   ├── tests/
│   │   ├── main.py
│   │   ├── pyproject.toml
│   │   └── Dockerfile
│   ├── frontend/             # Frontend container build context
│   ├── icons/                # Unraid Docker icons
│   └── docker-compose.yml
├── docker/
│   └── unraid/               # Single-container Unraid image build files
├── templates/
│   └── scenewhisper.xml      # Community Applications Docker template
├── ca_profile.xml            # Community Applications repository profile
└── README.md

Prerequisites

  • Docker
  • Docker Compose v2
  • NVIDIA drivers and NVIDIA Container Toolkit for GPU acceleration
  • uv for local backend development and tests

Check the host tools:

docker --version
docker compose version
nvidia-smi

Docker Compose

Clone the repository:

git clone git@github.com:Master-Oogwayxxx/SceneWhisper.git
cd SceneWhisper

Run Compose commands from the app directory:

cd app
docker compose build
docker compose up -d

The default services are:

  • SceneWhisper-worker: backend API and worker on http://localhost:8010
  • SceneWhisper: frontend UI on http://localhost:5175

Stop the stack:

docker compose down

Follow logs:

docker compose logs -f backend
docker compose logs -f frontend

Volumes and Ports

The Compose file uses these defaults:

Host path Container path Purpose
/mnt/user/videos /videos Video library
/mnt/user/appdata/docker-compose-projects/SceneWhisper/data /data Runtime app data
/mnt/user/appdata/docker-compose-projects/SceneWhisper/hf-cache /root/.cache/huggingface Hugging Face model cache

Default ports:

  • Frontend: 5175
  • Backend: 8010

GPU and CPU Hosts

GPU hosts use all available NVIDIA GPUs by default:

NVIDIA_VISIBLE_DEVICES: all
NVIDIA_DRIVER_CAPABILITIES: compute,utility,video

For CPU-only hosts, remove or comment the Compose GPU device reservation and set:

NVIDIA_VISIBLE_DEVICES: void

You can verify Docker GPU access with:

docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi

Backend Development

Run backend commands from app/backend:

uv run pytest
uv run uvicorn main:app --reload --host 0.0.0.0 --port 8010

Backend source is under app/backend/src, with tests in app/backend/tests.

Unraid Community Applications

SceneWhisper includes Unraid Community Applications metadata and a dedicated single-container image definition:

ca_profile.xml
templates/scenewhisper.xml
docker/unraid/Dockerfile

The published image is expected at:

ghcr.io/master-oogwayxxx/scenewhisper:latest

Template defaults:

  • Web UI: http://[IP]:5175
  • Video library: /mnt/user/videos mapped to /videos
  • App data: /mnt/user/appdata/scenewhisper/data mapped to /data
  • Model cache: /mnt/user/appdata/scenewhisper/hf-cache
  • Optional secret: HF_TOKEN

Unraid GPU users should install the NVIDIA Driver plugin and confirm Docker GPU runtime support is available. CPU-only users can set NVIDIA_VISIBLE_DEVICES to void in the template advanced settings.

For local template testing before Community Applications submission, add this template repository URL in Unraid Docker settings:

https://github.com/Master-Oogwayxxx/SceneWhisper/tree/main/templates

Maintainer Release Flow

  1. Push changes to main or run the Publish Unraid Image workflow manually.
  2. Confirm the GHCR package ghcr.io/master-oogwayxxx/scenewhisper:latest is public.
  3. Validate the repository at https://ca.unraid.net/submit/new.
  4. Submit the public GitHub repository URL after validation passes.

Before submitting to Community Applications, confirm:

  • .github/workflows/publish-unraid-image.yml is committed so GHCR publishes on main.
  • ghcr.io/master-oogwayxxx/scenewhisper:latest exists and is public.
  • ca_profile.xml and templates/scenewhisper.xml use reachable raw GitHub URLs from the public repository.
  • The icon URL resolves to app/icons/scene_whisper.png.

Troubleshooting

If the GPU is not detected, confirm the NVIDIA Container Toolkit is installed and Docker has been restarted:

sudo apt install nvidia-container-toolkit
sudo systemctl restart docker

If a port is already in use, check the conflicting process or change the port mapping in app/docker-compose.yml:

lsof -i :8010
lsof -i :5175

License

Apache License 2.0. See LICENSE.

Requirements

NVIDIA GPU users must install the NVIDIA Driver plugin and NVIDIA Container Toolkit support. CPU-only operation can set NVIDIA_VISIBLE_DEVICES to void.

Related apps

Details

Repository
ghcr.io/master-oogwayxxx/scenewhisper:latest
Last Updated2026-10-11
First Seen2026-10-11

Runtime arguments

Web UI
http://[IP]:[PORT:80]/
Network
bridge
Privileged
false

Template configuration

WebUI PortPorttcp

SceneWhisper web interface.

Target
80
Default
5175
Value
5175
Video LibraryPathrw

Host path containing videos to scan and process.

Target
/videos
Default
/mnt/user/videos
Value
/mnt/user/videos
App DataPathrw

Database, logs, and runtime state.

Target
/data
Default
/mnt/user/appdata/scenewhisper/data
Value
/mnt/user/appdata/scenewhisper/data
Hugging Face CachePathrw

Model cache for Whisper and related downloads.

Target
/root/.cache/huggingface
Default
/mnt/user/appdata/scenewhisper/hf-cache
Value
/mnt/user/appdata/scenewhisper/hf-cache
TimezoneVariable

Container timezone.

Target
TZ
Default
America/Los_Angeles
Value
America/Los_Angeles
Hugging Face TokenVariable

Optional token for private or rate-limited Hugging Face model access.

Target
HF_TOKEN
NVIDIA Visible DevicesVariable

Use all NVIDIA GPUs, a specific GPU id, or void for CPU-only mode.

Target
NVIDIA_VISIBLE_DEVICES
Default
all
Value
all
NVIDIA Driver CapabilitiesVariable

NVIDIA runtime capabilities used by AI and ffmpeg workloads.

Target
NVIDIA_DRIVER_CAPABILITIES
Default
compute,utility,video
Value
compute,utility,video
Hugging Face TransferVariable

Enable faster Hugging Face model downloads.

Target
HF_HUB_ENABLE_HF_TRANSFER
Default
1
Value
1