10 Commits
Author SHA1 Message Date
root 95bb7c50ee feat: point the fleet at wss LiveKit room uwh-telhai
Publish and subscribe over wss://livekit.uni-wh.de:7800 and refuse
cleartext ws://. Conference room is uwh-telhai. Includes the uncommitted
encoded H.264 publish path, Rally hairpin, and KMS wall overlay.
2026-10-11 00:48:50 +00:00
root ef9573c294 feat: OpenVINO CPU MetricGAN mask; 100 ms hop
Per-worker OpenVINO CPU runs the MetricGAN BLSTM mask; SpeechBrain
keeps STFT/ISTFT. ENHANCE_INFER=auto with torch fallback. Hop is
0.1 s after isolated last_ms ~6 ms (model+ov+beamform).
2026-08-28 07:25:30 +00:00
root 42443420be chore: OpenVINO MetricGAN GPU spike (do not merge)
Arc B580 lost the kill test: batch-14 mask p50 103 ms vs Torch CPU
70 ms enhance / OpenVINO CPU 5 ms. Production stays on known-good
CPU workers (tag known-good-cpu-2026-08-28). Hop stays 200 ms.
2026-08-28 07:18:54 +00:00
root 736eccb932 feat: known-good CPU audio + 1080p VAAPI + KMS display wall
Working snapshot before the OpenVINO MetricGAN GPU spike:
thread-capped SpeechBrain DelaySum+MetricGAN on CPU, hop 200 ms,
video hold for A/V sync, Arc VAAPI H.264, headless kmssink wall.
32 tests passing. Do not treat this as a GPU audio path.
2026-08-28 07:14:21 +00:00
root 11776adf52 chore: Arc B580 offload spike script
Read-only probe: CUDA unusable, OpenCL sees the B580, OpenVINO not
installed. Per-worker GPU still out of scope.
2026-08-28 06:26:02 +00:00
root f83f0f0683 docs: latency budget, thread cap, and Arc B580 CUDA mismatch
Also extend bench_enhance.py with --stereo --hop for isolated RTF.
2026-08-28 06:21:02 +00:00
root 36c4bbc956 perf: ALSA ffmpeg low-delay flags
nobuffer, low_delay, and a small thread_queue_size on the camera
mic capture process.
2026-08-28 06:19:44 +00:00
root b02138669a perf: reuse GCC-PHAT TDOA; skip 16 kHz AudioNormalizer
Delay-and-Sum still uses the tutorial STFT/cov/GCC-PHAT/DelaySum/ISTFT
path. TDOA is refreshed every BEAMFORM_TDOA_EVERY hops (default 5).
2026-08-28 06:19:02 +00:00
root 23c2e9c4c0 perf: cap torch threads per camera worker
Default TORCH_NUM_THREADS=1 so 14-20 workers do not each spawn a
28-thread BLAS pool.
2026-08-28 06:17:33 +00:00
root 950128deec feat: DelaySum beamform plus hop process latency
Keep C920 stereo through SpeechBrain GCC-PHAT Delay-and-Sum, then
MetricGAN, and publish mono to LiveKit. Record last_dt_ms and log
process_ms/capture_ms every 50 hops for the latency plan.
2026-08-28 06:16:25 +00:00