Publish and subscribe over wss://livekit.uni-wh.de:7800 and refuse cleartext ws://. Conference room is uwh-telhai. Includes the uncommitted encoded H.264 publish path, Rally hairpin, and KMS wall overlay.
livekit-cameras
Publishes 20 Logitech C920 webcam streams (video + cleaned audio) to a LiveKit room.
Each camera becomes one LiveKit participant (cam-01 .. cam-20) that
publishes:
- a camera video track (640x360 @ 30 fps MJPG, H.264 VAAPI on Arc). 1280x720@30 and 1920x1080@30 negotiate, but UVC on the 6–7 camera USB 2.0 hubs fails buffer-pool allocation for most devices. Keep 30 fps; drop resolution if USB isoc is exhausted. YUYV 1080p is 5 fps only — always MJPEG.
- a microphone audio track (16 kHz mono, cleaned with one shared speechbrain MetricGAN process; workers do not each load the model)
Identity is the USB serial, pinned by udev (/dev/camNN + slots.json).
--cameras 1,5 means participants cam-01 and cam-05, not discovery order.
Layout
run.py orchestrator: discovery + per-camera worker processes
run_worker.py one camera: ffmpeg capture -> LiveKit track publish
discovery.py serial-based camera <-> ALSA card discovery
audio_cleanup.py speechbrain SpectralMaskEnhancement (16 kHz) + fallback
tokens.py LiveKit JWT generation
config.py .env / environment configuration
prefetch_model.py one-time speechbrain model download
generate_udev_rules.py (re)generate udev pinning + slots.json from live state
slots.json stable slot map: cam-NN -> USB serial, /dev/camNN, ALSA card
smoke_test.py one-camera LiveKit publish with enhancement ON
scripts/bench_enhance.py RTF bench (1 s window / 200 ms hop)
deploy/livekit-cameras.service optional systemd unit (do not enable blindly)
Setup
uv venv --python 3.12
uv pip install -p .venv/bin/python -r pyproject.toml # or: uv sync
cp .env.example .env # then fill in LIVEKIT_URL / API_KEY / API_SECRET
.venv/bin/python prefetch_model.py # download speechbrain model once
Stable camera identity (cam-01 is always cam-01)
Each C920 has a unique built-in USB serial — that is the camera's UUID.
udev rules (/etc/udev/rules.d/99-c920-pin.rules, generated from the live
state) bind every slot to a serial, so identity survives reboots, hub
reordering, and any connect sequence:
- video:
/dev/cam01../dev/cam20(primary UVC node;.meta= metadata) - audio: fixed ALSA card per slot,
hw:2,0..hw:21,0(SOUND_CARD_INDEX)
discovery.py assigns identities from slots.json (serial lookup, with a
port-based fallback for cameras not yet in the map). To re-map slots (e.g.
after physically swapping cameras), regenerate:
.venv/bin/python generate_udev_rules.py --write
sudo udevadm control --reload-rules
sudo udevadm trigger --subsystem-match=video4linux --action=add
sudo udevadm trigger --subsystem-match=sound --action=add
The generator bakes the current physical wiring as canonical. Do not
re-run --write unless you intend to remap slots.
Run
.venv/bin/python run.py # all 20 cameras
.venv/bin/python run.py --cameras 1,5 # subset
.venv/bin/python run.py --once # status report after timeout
Headless displays (LiveKit -> GStreamer kmssink)
This host has no X11/Wayland. run_displays.py joins the LiveKit room as
display-wall (subscribe-only) and drives three DRM roles with GStreamer
kmssink:
- grid — 5x4 mosaic of cam-01..cam-20
- speaker — designated speaker camera (
DISPLAY_SPEAKER_CAMERA, defaultcam-01; set toactiveto follow LiveKit active speaker) - screenshare — stub placeholder until the production LiveKit server is online (no subscribe, no speaker login yet)
Default connector order: connected Arc HDMI first (today HDMI-A-7 is live; HDMI-A-5 / HDMI-A-6 light up when cabled).
.venv/bin/python run_displays.py --list
.venv/bin/python run_displays.py --test-pattern
.venv/bin/python run_displays.py
.venv/bin/python run_displays.py --speaker-camera cam-07
sudo systemctl restart livekit-displays
journalctl -u livekit-displays -f
.env: DISPLAY_ROLES, DISPLAY_CONNECTORS, DISPLAY_SPEAKER_CAMERA,
DISPLAY_GRID_*. Local cameras publish LiveKit attributes tag=uwh
(PARTICIPANT_TAGS). The wall can hide those with DISPLAY_HIDE_TAGS=uwh
so remotes fill the mosaic instead; leave DISPLAY_HIDE_TAGS empty to keep
showing the local fleet (current default). kmssink takes the DRM master (framebuffer console on
HDMI-A-7 is replaced by the grid).
Audio cleanup
C920 mics are 16 kHz stereo. Each hop (default 100 ms) is kept stereo through a SpeechBrain preprocessor, then published as mono:
- Delay-and-Sum beamform (GCC-PHAT TDOA, frozen for
BEAMFORM_TDOA_EVERY=5hops) — live path from the SpeechBrain multi-microphone beamforming tutorial speechbrain/metricgan-plus-voicebankneural enhancement on a 1 s rolling window (ENHANCE_CHUNK_S=1.0,ENHANCE_HOP_S=0.1). The MetricGAN mask runs on OpenVINO CPU (ENHANCE_INFER=auto); STFT/ISTFT stay in Torch.ENHANCE_INFER=torchforces SpeechBrainenhance_batch. Do not use the Arc B580 for this net.
The 1 s window is past context, not +1 s delay: only the latest hop is published. Algorithmic delay ≈ hop + compute.
Optional latency flags (all default to current behavior). Suggested wall
A/B: ENHANCE_HOP_S=0.05 AUDIO_QUEUE_MS=60 WORKER_SPLIT_EXECUTOR=1 ENHANCE_SPEAKER_ONLY=1 DISPLAY_VIDEO_CAPACITY=1 DISPLAY_VIDEO_FORMAT=bgra DISPLAY_GST_QUEUE_BUFFERS=1 DISPLAY_SPEAKER_PUSH=arrival. Aggressive:
ENHANCE_MODE=off BEAMFORM_MODE=off ENHANCE_HOP_S=0.02 AUDIO_QUEUE_MS=40 VIDEO_HOLD_S=0. See .env.example.
Each worker caps Torch/BLAS/OpenVINO to TORCH_NUM_THREADS=1 (default)
so 14–20 processes do not oversubscribe the i7-14700. This host has an
Intel Arc B580; the venv Torch is NVIDIA CUDA and cannot use it. Do
not set run_opts={"device":"cuda"}.
One worker process per camera loads the modules once and reuses them.
enhance_batch must be called with lengths=torch.tensor([1.0])
(relative, not sample count). ENHANCE_MODE=off (or a failed model
load with auto) skips MetricGAN; BEAMFORM_MODE=off averages L/R
instead of Delay-and-Sum. DSP high-pass + soft-clip is the last-resort
fallback so streams still work.
Until the 1 s window is full, hops use the DSP fallback so LiveKit is
not silent on startup. Workers log process_ms / capture_ms every 50
hops.
Appliance boot / power button
This box is meant to behave like an appliance:
- Power button while off → BIOS/OS boot →
livekit-cameraswaits for USB cameras + LiveKit, then publishes. - Power button while on → full
poweroff(not sleep). Suspend/hibernate targets are masked. Drop-in:/etc/systemd/logind.conf.d/50-power-button.conf
systemd (enabled)
livekit-server.service— local TEST LiveKit on :7880 (disabled; cameras no longerWants=it). Do not re-enable unless falling back to localhost.livekit-cameras.service— publishers. Destination isLIVEKIT_URLin.env(wss://livekit.uni-wh.de:7800, roomuwh-telhai). Signaling is WSS only; the client verifies the TLS certificate.livekit-displays.service— optional. Subscribes to the camera room and drives 3 DRM connectors with GStreamerkmssink(no X11/Wayland).
systemctl status livekit-cameras livekit-server
journalctl -u livekit-cameras -f
systemctl restart livekit-cameras # pick up newly plugged cameras
Notes
- Video uses MJPG directly from the UVC interface (
ffmpeg -f v4l2 -input_format mjpeg) — cheap, and converts to BGRA for LiveKit. - Audio: ALSA mics on these C920s deliver 16 kHz stereo only; the worker beamforms (Delay-and-Sum) then enhances, and publishes mono to LiveKit.
- Camera identity is USB serial via
slots.json/ udev; USB port is fallback only for cameras not yet in the map. - Workers auto-restart (up to 10 times) if a camera drops.
- Local smoke:
.venv/bin/python smoke_test.py(needs livekit-server on 7880).