The TV stays on
How we copied the DVR feed to the browser without blacking out the store screen
In short: I copied the DVR video into the internal portal without touching the strip that goes to the store TV. A tee splits the signal, the encoder runs in its own process, and the always-on memory copies went from 249 MB/s to 62 MB/s.
Every time I touched the DVR video capture, the store's vertical TV ended up black in front of customers. That screen shows a live strip from the street camera: the image comes out of the DVR's HDMI (the recorder for the four cameras), enters through a Blackmagic DeckLink capture card, and a GStreamer process on the kiosk machine crops it. On top of that, I wanted the full mosaic, the four cameras together, in the internal portal.
Why the TV branch comes first
Whoever walks by on the sidewalk looks at the strip. The mosaic in the portal saves me the walk to the warehouse to check the DVR screen. That order is written into the pipeline, the chain of steps the video travels through: the portal branch drops frames when it falls behind, and the TV branch never drops any.
Analogy: The guard watching the camera monitor shouldn't also be narrating what he sees to the boss over the phone, because it pulls him off his real job. You put a second screen next to him with the same feed, and someone else turns on a phone camera only when the boss calls.
A tee right after the capture
I put a tee (a fork in the signal) right after decklinkvideosrc, the element that reads the card. Branch A crops the strip and sends it to the TV, same as before. Branch B publishes the full 1920x1080 frame to a shared memory socket, a mailbox in RAM that another process can read. That branch sits behind a "leaky" queue: when it fills up, it throws out the old frames instead of making everyone else wait. A viewer stuck in the browser no longer holds back the TV branch.
The NVENC encoder compresses that video with the GPU and it's the most expensive part of the system, so it runs in a separate process. go2rtc, the server that hands the video to the browser, starts it only when someone opens the stream in the portal.
The encoder inside left the strip black
My first version was the obvious one and I got it wrong: I put NVENC inside the same capture process. Creating the CUDA context (the work session with the GPU) blocks the pipeline startup. That delayed the card's initial handshake, which negotiates format and frequency with the HDMI before it passes a single frame, and the strip went black with the not-negotiated error. I reverted it the same day, and the rule stuck: the encoder never shares a process with the capture.
Then two silent failures showed up, the kind that gives me the most work. shmsink, the element that writes to shared memory, fails without raising an error if it finds an old socket from the previous startup, so the portal stream was dead while the TV looked perfect. Now the script deletes the socket before starting. The other one: branch B was copying 60 frames per second that nobody watched, because the dropping happened after the copy. I moved the dropping to the producer and the branch settled at 15 frames per second (fps).
What changed in numbers
The portal shows the mosaic at 1080p and 15 fps. The always-on branch went from 249 MB/s to 62 MB/s of memory copies, and the capture process went from between 46% and 52% of CPU to 32%. Opening the stream from the portal doesn't make that idle cost worse, because the expensive process runs apart and only while someone is watching.
The watchdog that checks every 20 seconds whether the TV strip froze doesn't watch the portal branch, so if that branch dies in silence again I'll find out when I open the browser. The portal viewer also sees jumps when their connection falls behind, and I asked for it that way.
What I do differently now
Three rules went into the document for the kiosk VM. Before I write a pipeline that shares a machine with something the customer sees, I decide which queue drops frames and how much CPU each part gets. Every new branch also needs its own sign of life, because we had all four screens black for 5 hours and 20 minutes while the system reported itself healthy. And the expensive part starts only when someone asks for it: I had already killed an object detection that ate a whole core competing with video playback, and today the encoder starts only when the boss calls.
More in Infrastructure
This is the only post here so far. Browse the section or all posts.