Files
root 95bb7c50ee feat: point the fleet at wss LiveKit room uwh-telhai
Publish and subscribe over wss://livekit.uni-wh.de:7800 and refuse
cleartext ws://. Conference room is uwh-telhai. Includes the uncommitted
encoded H.264 publish path, Rally hairpin, and KMS wall overlay.
2026-10-11 00:48:50 +00:00

7.9 KiB
Raw Permalink Blame History

livekit-cameras

Publishes 20 Logitech C920 webcam streams (video + cleaned audio) to a LiveKit room.

Each camera becomes one LiveKit participant (cam-01 .. cam-20) that publishes:

  • a camera video track (640x360 @ 30 fps MJPG, H.264 VAAPI on Arc). 1280x720@30 and 1920x1080@30 negotiate, but UVC on the 6–7 camera USB 2.0 hubs fails buffer-pool allocation for most devices. Keep 30 fps; drop resolution if USB isoc is exhausted. YUYV 1080p is 5 fps only — always MJPEG.
  • a microphone audio track (16 kHz mono, cleaned with one shared speechbrain MetricGAN process; workers do not each load the model)

Identity is the USB serial, pinned by udev (/dev/camNN + slots.json). --cameras 1,5 means participants cam-01 and cam-05, not discovery order.

Layout

run.py              orchestrator: discovery + per-camera worker processes
run_worker.py       one camera: ffmpeg capture -> LiveKit track publish
discovery.py        serial-based camera <-> ALSA card discovery
audio_cleanup.py    speechbrain SpectralMaskEnhancement (16 kHz) + fallback
tokens.py           LiveKit JWT generation
config.py           .env / environment configuration
prefetch_model.py   one-time speechbrain model download
generate_udev_rules.py  (re)generate udev pinning + slots.json from live state
slots.json          stable slot map: cam-NN -> USB serial, /dev/camNN, ALSA card
smoke_test.py       one-camera LiveKit publish with enhancement ON
scripts/bench_enhance.py  RTF bench (1 s window / 200 ms hop)
deploy/livekit-cameras.service  optional systemd unit (do not enable blindly)

Setup

uv venv --python 3.12
uv pip install -p .venv/bin/python -r pyproject.toml   # or: uv sync
cp .env.example .env    # then fill in LIVEKIT_URL / API_KEY / API_SECRET
.venv/bin/python prefetch_model.py     # download speechbrain model once

Stable camera identity (cam-01 is always cam-01)

Each C920 has a unique built-in USB serial — that is the camera's UUID. udev rules (/etc/udev/rules.d/99-c920-pin.rules, generated from the live state) bind every slot to a serial, so identity survives reboots, hub reordering, and any connect sequence:

  • video: /dev/cam01 .. /dev/cam20 (primary UVC node; .meta = metadata)
  • audio: fixed ALSA card per slot, hw:2,0 .. hw:21,0 (SOUND_CARD_INDEX)

discovery.py assigns identities from slots.json (serial lookup, with a port-based fallback for cameras not yet in the map). To re-map slots (e.g. after physically swapping cameras), regenerate:

.venv/bin/python generate_udev_rules.py --write
sudo udevadm control --reload-rules
sudo udevadm trigger --subsystem-match=video4linux --action=add
sudo udevadm trigger --subsystem-match=sound --action=add

The generator bakes the current physical wiring as canonical. Do not re-run --write unless you intend to remap slots.

Run

.venv/bin/python run.py                  # all 20 cameras
.venv/bin/python run.py --cameras 1,5    # subset
.venv/bin/python run.py --once           # status report after timeout

Headless displays (LiveKit -> GStreamer kmssink)

This host has no X11/Wayland. run_displays.py joins the LiveKit room as display-wall (subscribe-only) and drives three DRM roles with GStreamer kmssink:

  1. grid — 5x4 mosaic of cam-01..cam-20
  2. speaker — designated speaker camera (DISPLAY_SPEAKER_CAMERA, default cam-01; set to active to follow LiveKit active speaker)
  3. screenshare — stub placeholder until the production LiveKit server is online (no subscribe, no speaker login yet)

Default connector order: connected Arc HDMI first (today HDMI-A-7 is live; HDMI-A-5 / HDMI-A-6 light up when cabled).

.venv/bin/python run_displays.py --list
.venv/bin/python run_displays.py --test-pattern
.venv/bin/python run_displays.py
.venv/bin/python run_displays.py --speaker-camera cam-07

sudo systemctl restart livekit-displays
journalctl -u livekit-displays -f

.env: DISPLAY_ROLES, DISPLAY_CONNECTORS, DISPLAY_SPEAKER_CAMERA, DISPLAY_GRID_*. Local cameras publish LiveKit attributes tag=uwh (PARTICIPANT_TAGS). The wall can hide those with DISPLAY_HIDE_TAGS=uwh so remotes fill the mosaic instead; leave DISPLAY_HIDE_TAGS empty to keep showing the local fleet (current default). kmssink takes the DRM master (framebuffer console on HDMI-A-7 is replaced by the grid).

Audio cleanup

C920 mics are 16 kHz stereo. Each hop (default 100 ms) is kept stereo through a SpeechBrain preprocessor, then published as mono:

  1. Delay-and-Sum beamform (GCC-PHAT TDOA, frozen for BEAMFORM_TDOA_EVERY=5 hops) — live path from the SpeechBrain multi-microphone beamforming tutorial
  2. speechbrain/metricgan-plus-voicebank neural enhancement on a 1 s rolling window (ENHANCE_CHUNK_S=1.0, ENHANCE_HOP_S=0.1). The MetricGAN mask runs on OpenVINO CPU (ENHANCE_INFER=auto); STFT/ISTFT stay in Torch. ENHANCE_INFER=torch forces SpeechBrain enhance_batch. Do not use the Arc B580 for this net.

The 1 s window is past context, not +1 s delay: only the latest hop is published. Algorithmic delay ≈ hop + compute.

Optional latency flags (all default to current behavior). Suggested wall A/B: ENHANCE_HOP_S=0.05 AUDIO_QUEUE_MS=60 WORKER_SPLIT_EXECUTOR=1 ENHANCE_SPEAKER_ONLY=1 DISPLAY_VIDEO_CAPACITY=1 DISPLAY_VIDEO_FORMAT=bgra DISPLAY_GST_QUEUE_BUFFERS=1 DISPLAY_SPEAKER_PUSH=arrival. Aggressive: ENHANCE_MODE=off BEAMFORM_MODE=off ENHANCE_HOP_S=0.02 AUDIO_QUEUE_MS=40 VIDEO_HOLD_S=0. See .env.example.

Each worker caps Torch/BLAS/OpenVINO to TORCH_NUM_THREADS=1 (default) so 14–20 processes do not oversubscribe the i7-14700. This host has an Intel Arc B580; the venv Torch is NVIDIA CUDA and cannot use it. Do not set run_opts={"device":"cuda"}.

One worker process per camera loads the modules once and reuses them. enhance_batch must be called with lengths=torch.tensor([1.0]) (relative, not sample count). ENHANCE_MODE=off (or a failed model load with auto) skips MetricGAN; BEAMFORM_MODE=off averages L/R instead of Delay-and-Sum. DSP high-pass + soft-clip is the last-resort fallback so streams still work.

Until the 1 s window is full, hops use the DSP fallback so LiveKit is not silent on startup. Workers log process_ms / capture_ms every 50 hops.

Appliance boot / power button

This box is meant to behave like an appliance:

  • Power button while off → BIOS/OS boot → livekit-cameras waits for USB cameras + LiveKit, then publishes.
  • Power button while on → full poweroff (not sleep). Suspend/hibernate targets are masked. Drop-in: /etc/systemd/logind.conf.d/50-power-button.conf

systemd (enabled)

  • livekit-server.service — local TEST LiveKit on :7880 (disabled; cameras no longer Wants= it). Do not re-enable unless falling back to localhost.
  • livekit-cameras.service — publishers. Destination is LIVEKIT_URL in .env (wss://livekit.uni-wh.de:7800, room uwh-telhai). Signaling is WSS only; the client verifies the TLS certificate.
  • livekit-displays.service — optional. Subscribes to the camera room and drives 3 DRM connectors with GStreamer kmssink (no X11/Wayland).
systemctl status livekit-cameras livekit-server
journalctl -u livekit-cameras -f
systemctl restart livekit-cameras   # pick up newly plugged cameras

Notes

  • Video uses MJPG directly from the UVC interface (ffmpeg -f v4l2 -input_format mjpeg) — cheap, and converts to BGRA for LiveKit.
  • Audio: ALSA mics on these C920s deliver 16 kHz stereo only; the worker beamforms (Delay-and-Sum) then enhances, and publishes mono to LiveKit.
  • Camera identity is USB serial via slots.json / udev; USB port is fallback only for cameras not yet in the map.
  • Workers auto-restart (up to 10 times) if a camera drops.
  • Local smoke: .venv/bin/python smoke_test.py (needs livekit-server on 7880).