Client Monitor · 4.7.0

Client Monitor 4.7.0

· balazskreith

Reworked score attribution and detectors, three new failure detectors, and ClientSample 3.6.0 migration changes.

View release on GitHub ↗

A large release, summarised by what actually changed rather than by the order it was built in. Three themes: scores now say who is responsible for what, detectors were rebuilt around the signal each one can actually see, and three new detectors cover failures nothing owned.

Breaking

  • sourceEncoderBottleneckDetector config is gone, replaced by outboundFrameSupplyDetector, encoderPerformanceDetector and inboundFrameSupplyDetector — one block per attachment point. Code that passed it (including : null to disable) must be updated; TypeScript flags it, JavaScript does not.
  • ClientEvent.payload, ClientIssue.payload, ClientMetaData.payload and ExtensionStat.payload are records on the wire, not pre-serialised JSON strings. Servers must read them as objects. Callers passing nested structures to addEvent/raiseIssue/addIssue/addMetaData must flatten or JSON-encode those values themselves — a nested structure in an event payload is now a compile error. (~12% escaping overhead removed from payload bytes, ~18% faster whole-sample serialisation.)
  • encode*ScoreReasons removed from the ScoreCalculator interface. Custom calculators implement update() and set reasons on the calculated scores; the sampling helper is exported as sampledScoreReasons.
  • low-bitrate-per-pixel and low-bitrate-for-resolution no longer exist as score reasons, along with BPP_RANGES, expectedVideoBitrate, VIDEO_BITRATE_EXPECTATION and calculateBaseVideoScore. pixelated-video replaces them.
  • high-packetloss and high-jitter no longer appear on track score reasons — only on the peer connection. See Score attribution below.
  • Schema ClientSample 3.6.0: scoreReasons is a Record<string, number> on the client, peer-connection and track entries.

Score attribution: the path is the connection’s problem, the picture is the track’s

The same network condition used to be charged up to three times. Measured over a captured session whose real streams ran at 0.00% loss and 2 ms jitter: mean client score 2.93 → 4.86, share of samples below 3.0 85% → 0%.

  • Streams carrying no media no longer measure the path. An SFU’s bandwidth-probation stream (mediasoup’s mid: "probator") delivers deliberately discardable packets and no frames — observed at ~2 kbps with ~50% “loss” and ~490 ms “jitter” beside real streams at 0% and 2 ms. Averaged in unweighted it produced high-jitter: 2 and high-packetloss: 2 on 99% of samples. A stream must now show evidence of carrying media before its ratios count: MIN_PATH_SAMPLE_BITRATE (8 kbps), MIN_PATH_SAMPLE_PACKETS (25/interval), or any frames. Any one suffices, so DTX audio and thin video still count. This alone moved the peer connection’s mean from 1.08 to 4.92.
  • Tracks are no longer charged for loss or jitter. Those are properties of the shared path and are already charged there. Charging them again could floor a track — an outbound audio track scored 0.03 on a path whose loss was 0% at the 95th percentile. Tracks keep every penalty measuring what the user perceived: freezes, low and volatile fps, dropped frames, pixelation, concealment, time-stretch, jitter-buffer delay. A server joins a track’s symptoms to its peer connection’s path reasons from the same sample.
  • The client score aggregation is unchanged: the peer connection still scales its tracks by pcScore / 5, so a degraded path degrades what the viewer got from every track on it and a call can never score better than the connection carrying it. Removing the track-level loss and jitter penalties is what stops the same packets being charged twice; the multiplication itself was only ever counting the path once.
  • RTT and jitter are separate peer-connection reasons — a long path and a jittery path are different problems (high-rtt at −1 above 150 ms, −2 above 300 ms; high-jitter at −1 above 30 ms, −2 above 100 ms). very-high-rtt is gone — it was the only reason key that split one condition across two keys instead of carrying the magnitude, which is what every other reason does. Loss uses the per-interval deltaFractionLost and is averaged across streams rather than summed.
  • Each entity ships only its own reasons. The client entry carried the aggregate of everything below it, so one track pixelating produced pixelated-video on the track entry and the client entry. ClientMonitor.scoreReasons now means what it means everywhere else — this entity’s own subtractions, of which there are none today. The aggregate moved to the 'score' event’s currentReasons for applications reacting live. sendScoreReasonsToServer (default true) drops reasons from the wire without touching scores.

Pixelation is measured from the quantizer, and charged by how big the picture is

  • pixelated-video reads the inbound qpSum. The old bitrate-per-pixel reason was wrong twice over: its floor ignored resolution (required bitrate scales as roughly pixels^0.75, so one floor cannot fit 180p and 1080p — in practice it demanded ~1 Mbps at 360p30 and took a loss-free 640×358@30 VP8 stream at 500 kbps to full saturation every sample), and dividing by measured fps meant a track halving its frame rate doubled its bits-per-pixel and shed the penalty — the metric rewarded dropping frames. QP is the encoder stating how coarsely it had to quantize, which is the blockiness the viewer is looking at. Where the browser reports no qpSum for the codec, the reason is absent entirely rather than modelled from bitrate. InboundRtpMonitor.avgQpPerFrame/deltaQpSum are new, mirroring the outbound side, and reset rather than carrying a stale average forward when nothing decoded.
  • Bands are per codec and per motion class (VIDEO_QP_THRESHOLDS, indexed [codec][motionType]). QP scales are not comparable as fractions of their ranges (H.264 0–51, VP8 0–127, VP9/AV1 0–255), and the same quantizer is not equally visible on all content — movement masks artifacts, a slide shows every blocked edge. Bands run the opposite way to bitrate. Undeclared, screen share is judged lowmotion and everything else standard. Unrecognised codec, no judgement. All ramps on DefaultScoreCalculator are mutable statics.
  • A large picture is charged harder than a small one, deliberately unfairly. The presented size selects the weight a saturated quantizer is worth: 3.0 when the picture is magnified ≥1.5× linear (PIXELATION_MAX_PENALTY_LARGE — a big video gone to blocks is the worst thing short of it stopping), 2.0 at roughly decoded size, 0.5 below 0.75× (a thumbnail nobody can see the blocks in). Magnification is sqrt(presentedArea / decodedArea), taken from the areas so a differently proportioned box is not magnification on width alone. Undeclared presented size means the ordinary 2.0 — nothing is substituted for a missing number.

Application-declared track context

  • ClientMonitor.setInboundTrackContext(trackId, ctx) / setOutboundTrackContext(trackId, ctx), and setContext() on either track monitor. Everything the application knows and the stats never reveal travels through one call per direction, taking a partial object. InboundTrackContext carries contentType, motionType, presentedResolution and videoTag; OutboundTrackContext carries contentType alone, because motion and presentation describe how a track is watched and the sender does not know.
  • Declarable before the track exists, in both directions: signaling often announces a guest’s screen share before a packet arrives. A declaration is applied immediately if the monitor exists and otherwise held pending, consumed by whichever peer connection first manifests the track. Separate maps per direction — one shared map meant a send and a receive track sharing an id could take each other’s declaration. No timers, no cleanup.
  • Contexts merge, they do not replace, on a live monitor and in the pending state alike, so a content type from signaling survives a later call that only attaches the video element. An explicit undefined means “not declared here”, not “reset”.
  • presentedResolution is declared in device pixels, or derived from videoTag and re-measured every tick, so going full-screen or resizing a panel is picked up. The derivation reads the element’s layout box (clientWidth/clientHeight × devicePixelRatio) with the frame’s aspect ratio fitted into it as object-fit: contain does — never videoWidth/videoHeight, which are the intrinsic decoded size and would make every magnification exactly 1. Applications using object-fit: cover should declare the resolution themselves.
  • One TrackContentType ('camera' | 'screenshare') in monitors/TrackMonitor.ts, replacing the identical InboundTrackContentType and OutboundTrackContentType. What differs between the directions is not the type but how it is arrived at: auto-detected from track.getSettings().displaySurface when sending, declared when receiving. track.contentHint is never used — applications set 'detail'/'text' on camera tracks too.
  • Screen share is now actually scored as screen share. The old special case tested contentHint !== 'screen', not a valid value, so it matched every track. Screen-share tracks skip the fps, bitrate-volatility and target-deviation penalties (meaningless on mostly-static VBR content) and are charged on downscaled-screenshare instead: encoded area below ½ of the captured surface −1, below ¼ −2, where shared text stops being readable. Inbound screen share skips low-fps/volatile-fps for the same reason.

Frame supply: four detectors in a 2×2, one job each

SourceEncoderBottleneckDetector is gone. In its place, each cell answers one question with the unit that question deserves:

frames going missing (averaged over a duration)the stage cannot keep up (consecutive ticks)
outboundcapture-bottleneck — OutboundFrameSupplyDetectorencoder-bottleneck — EncoderPerformanceDetector
inbounddecoder-bottleneck — InboundFrameSupplyDetectorvideo-decoder-overloaded — DecoderPerformanceDetector
  • Ticks and durations are not two spellings of the same thing. A tick count is a confidence floor — two independent stats reads agreed — and every encoder/decoder signal is a per-interval ratio a single read can get wrong. A duration is a persistence bar: the capture device stayed short long enough to matter. So minConsecutiveTicks: 2 on the performance detectors, durationInMs: 15_000 on the supply detectors. This is also why JitterBufferStressDetector, DecoderPerformanceDetector and StuckDecoderDetector keep tick counts.
  • capture-bottleneck averages rather than thresholding each tick. Add up the frames the source delivered and the time it had; when 15 s has accumulated, compare the average against getSettings().frameRate at captureFpsRatioThreshold (0.9). Two running totals, no history buffer. Averaging is what catches a camera degrading in bursts — 150 frames per 5 s tick becomes 132, back to 150, then 97, so most individual ticks look fine while the 15 s average reads 24.5 fps against a configured 30 — and it weights how far the source fell short, not merely how often. On the captured failure it raises at t=45 s, 30 seconds before the camera stopped; the previous consecutive-tick rule never reached its threshold at all, because the starving ticks were interleaved with healthy ones.
  • New issue decoder-bottleneck: frames arrived and the decoder did not turn enough of them into pictures. Deliberately narrow — the bar is the arrival rate, so a stream throttled to 5 fps that decodes cleanly is silent, and frames that never arrived remain the network’s story.
  • Capture and encoder are chained and mutually exclusive. If the capture device is short, that is the whole answer and the encoder is not judged: an encoder handed too few frames has nothing to answer for. EncoderPerformanceDetector reads isIssueActive('capture-bottleneck-track-<id>') rather than sharing a private field, so the dependency is inspectable — a quiet encoder is explained by an issue anyone can see. At most one of the two is ever active.
  • No baseline, no judgement. Nothing is substituted for a missing getSettings().frameRate: without a stated rate there is nothing to fall short of. The rate is always the frame counter differenced against measured elapsed time, never mediaSource.framesPerSecond — the browser’s own figure is coarse and smooths this exact stutter away, reporting 30 across an interval that delivered 132 frames in five seconds. MediaSourceMonitor.sourceFps is undefined, not 0, on a counter reset, so replaceTrack no longer presents a healthy new camera as a dead one.
  • Four situations are no longer judged, because a low frame rate in them is legitimate: screen shares (content-driven — declare a moving surface with setOutboundTrackContext(id, { contentType: 'camera' })), backgrounded tabs, paused or non-live senders, and collection gaps/settings changes/counter resets, which discard the window rather than interpret it. The gap threshold is derived from collectingPeriodInMs (maxTickGapInMs), not configured — a fixed millisecond value means something different at every period.
  • cpuLimitationShareThreshold defaults to null on the encoder detector: the browser’s CPU-limitation share is ignored unless configured. CpuPerformanceDetector already reports it, and the useful thing to do with the two is correlate them — which only works while encoder-bottleneck is derived without reading the same signal. Set 0.3 for the old behaviour.

Threshold caveat. 0.9 over 15 s comes from two captured sessions — one failure, one control, one user, one camera model. They catch that failure and stay silent on that control, and are otherwise unvalidated. Treat capture-bottleneck as observation-grade until a corpus sets the numbers; low light is the case most likely to trip them, since many webcams settle at 15 fps in a dim room while getSettings().frameRate still reports 30.

Pause and visibility: absence is not evidence

  • ClientMonitor.activeTab and the tab-visibility watcher (watchTabVisibility, on by default). Defaults to true and stays true when the watcher is disabled or no document exists, so false always means the tab really is hidden. Each transition is a TAB_VISIBILITY_CHANGED client event. Browsers throttle background tabs, so the detectors whose signals that corrupts stand down: CpuPerformanceDetector, DecoderPerformanceDetector, StuckDecoderDetector, PlayoutDiscrepancyDetector and FreezedVideoTrackDetector.
  • Paused producers and consumers no longer look like failures. MediasoupTransportBinding mirrors pause state onto the track monitors, keeping the two kinds of silence distinct: a producer’s pause lands on OutboundTrackMonitor.paused (the sender stopped for everyone), a consumer’s on InboundTrackMonitor.paused (this leg opted out; the producer may still feed everyone else). remoteOutboundTrackPaused keeps its meaning and stays application-set. Both are synced on the observer events and re-synced every tick, because track monitors are created lazily from stats — an event-only sync would lose a pause that happened before the monitor existed.
  • Ten detectors stand down on a paused track: both dry-track detectors, StuckDecoderDetector, AudioConcealmentDetector, JitterBufferStressDetector, plus CaptureFailureDetector (silence check only), FreezedVideoTrackDetector, AudioDesyncDetector, PlayoutDiscrepancyDetector and SimulcastLayerDetector.
  • Monotonic counters are swallowed, not merely skipped. Freeze counts and NetEQ correction counters keep climbing through a pause, so skipping the paused ticks would deliver the whole pause as one delta on the first tick back — a false alarm exactly when the media was restored. Observed in a captured session: a resuming consumer showed 2.7 M concealed samples in one interval and was healthy on the next tick.
  • Pause stands down the silence check only, on capture. capture-track-ended and capture-track-muted still fire while paused: a camera unplugged or seized by another application is a fact about the device, true whether or not anyone was receiving it, and an application resuming onto a device that has since disappeared needs to know.
  • Known gap: SynthesizedSamplesDetector reads MediaPlayoutMonitor, which has no link to a track monitor, so it cannot see a pause. Chrome’s media-playout stats are per output device rather than per track, so the association is not merely missing but ambiguous. Left as a documented gap rather than guessed at.

New detectors

  • BlockedTransportDetector raises blocked-transport on the firewall signature every other detector structurally misses: STUN keeps answering — the pair is succeeded, consent passes, iceConnectionState is connected — yet media does not traverse. STUN consent responses count into the pair’s bytesReceived, so it never looks dry. Requires three sustained pieces of evidence over thresholdInMs (5 s): STUN demonstrably alive, the application demonstrably producing, and the media demonstrably not traversing. The payload’s evidence separates media-not-leaving-transport (host firewall, blocked socket, dead route) from no-return-traffic (DPI / UDP-throttling middlebox). Sending side only, where the client holds both halves of the proof.
  • NoAvailableIceCandidateDetector raises no-available-ice-candidate when the client cannot even begin: ICE gathering produced zero local candidates while the connection falls to disconnected/failed (immediately) or sits past thresholdInMs (6 s). Every other ICE issue describes a path that existed and stopped working; this one says no path was ever possible — no interface, airplane mode, a VPN that tore down every route. Never fires on a connection that once reached connected.
  • MediaPipelineDetector raises media-pipeline-stalled, the stage classifier: every pipeline stage has a monotonic counter proving progress, and a disruption is the first boundary where the upstream counter advances and the downstream one does not. stage: 'rtp-sender' — frames encode while no packet leaves. stage: 'transport-demux' — the ICE transport receives well above what RTCP and STUN can explain while every inbound RTP stays flat. suspectedIssueTypes links the specialist issues active at raise time, so one entry both localises the first broken stage and points at the evidence; registered last among the peer-connection detectors for that reason.
  • Supporting: PeerConnectionMonitor.iceGatheringState, IceCandidateMonitor.direction, PeerConnectionMonitor.localIceCandidates.

Detector correctness pass

A sweep of all 28 detectors for machinery that buys nothing and for fabricated baselines. Seven of these change a verdict; each is a case that was silently wrong.

  • ?? 0 on a deciding input is a fabricated baseline, and three of them were suppressing real detections. DecoderPerformanceDetector treated a missing deltaFractionLost as a perfectly quiet network and handed the decoder the blame for exactly the misattribution the check exists to prevent. AudioConcealmentDetector treated a missing deltaSilentConcealedSamples as zero, counting every silent moment as audible damage and defeating the one subtraction the detector is built around. StuckDecoderDetector treated a missing bitrate as a dead pipe, tearing down the accumulating stretch every tick — on any adapter that omits the field the detector could never fire at all. All three now hold their judgement instead.
  • StuckDecoderDetector could never fire without pliCount either, for the same reason; and its minStuckTicks config is removed — the time threshold already guarantees at least two ticks, so at the shipped default it could not change any verdict, and above it it was a second persistence bar in a different unit.
  • PlayoutDiscrepancyDetector silently skipped its own maximal case. A truthiness guard on deltaFramesRendered meant “everything arrived and nothing reached the screen” was never judged. ewmaFps also vetoed detection despite appearing only in the payload; it is now optional there and gates nothing.
  • EncoderPerformanceDetector left encoder-bottleneck open forever when the highest simulcast layer deactivated — a bare return where every other exit stands down.
  • CongestionDetector’s low sensitivity required an RTT reading it never uses, so a bandwidth-limited connection losing >5% was silently not congested before the first RTCP receiver report.
  • SynthesizedSamplesDetector never emitted its client event. A truthy createEvent test against a default that left the field unset made EXCESSIVE_SYNTHESIZED_AUDIO unreachable in every default build. It now spells the check === false like every sibling, and the default sets createEvent: true.
  • AudioDesyncDetector was comparing cumulative counters against zero. Its _prevCorrectedSamples was initialised to 0 and never updated, so what it called a per-tick delta was the whole session’s correction total — a fraction that could only climb. It now reads the deltas InboundRtpMonitor already computes, which also removes the duplicate subtraction.

Same-verdict simplifications: payload-only running totals and their sliding-window bookkeeping removed from FreezedVideoTrackDetector (firRate, keyFrameRate) and AudioConcealmentDetector (concealmentEventRate, burstiness); MAX_WINDOW_ENTRIES caps that were unreachable at any sane collecting period; _xxxOn booleans paired with _xxxStartedAt timestamps collapsed to the timestamp alone in five detectors; a duplicated ended-payload, a re-tested condition, and two spellings of one assignment.

CPU performance detector: frame-arrival burst guard

  • Bursty frame arrival no longer reads as CPU limitation. A simulcast layer switch, keyframe recovery or post-stall queue flush delivers a pile of frames inside one collect interval; the decoder trails that spike for exactly that interval and the decoded/received ratio dips while the machine is idle — observed as repeated one-tick dips on an idle machine, each coinciding with a VIDEO_RESOLUTION_CHANGED and recovering to ~1.0 next tick. The detector now keeps a smoothed (EWMA, α = 0.3) arrival baseline per inbound video ssrc and skips the ratio judgement on any interval exceeding frameArrivalBurstFactor × baseline (default 2.5), and on a track’s first interval. Sustained starvation still alerts, because its low ratio persists across ordinary-arrival intervals. Set the factor to undefined for the old behaviour.

Replay harness CLI

  • npm run replay -- <stats.jsonl> runs a captured session through a real ClientMonitor on a virtual clock and prints every detector fire as NDJSON on stdout — one record per fire with the issue’s own timestamp plus the tick/tickTimestamp locating it in the input, then a summary with per-type counts. Warnings go to stderr, so output pipes into jq or a corpus runner sweeping thresholds. --only, --config, --pretty, --updates, --real-time, --no-summary; reads stdin with -. Note a per-detector config block replaces the shipped one wholesale, so a sweep must pass the whole block. No new dependency: scripts/replay.mjs compiles to CommonJS under .replay/. tests/fixtures/degrading-camera.jsonl replays the interleaved-degradation shape through the real stats path.

Per-detector issue sampling

  • Every issue-raising detector exposes includeIssueInSample = true next to disabled. Set false and the detector keeps working locally — events fire, activeIssues and the issue lifecycle are maintained — but neither the raise nor the resolution is buffered into the ClientSample. raiseIssue/addIssue accept includeInSample directly. Useful for shrinking the sample when an issue is derivable server-side; the per-issue derivability table is in docs/DETECTOR_SAMPLE_WORTHINESS.md.

Other

  • Data channels carry their mediasoup identity: the binding stamps dataProducerId/dataConsumerId onto the matching DataChannelMonitor.attachments, joining sctpStreamParameters.streamId against RTCDataChannelStats.dataChannelIdentifier, retried each tick until the lazily-created monitor appears.
  • Fixed: the ICE_CANDIDATE event payload was built by spreading an RTCIceCandidate, whose properties are prototype getters — the spread copied nothing.
  • New monitor events 'blocked-transport', 'no-available-ice-candidate', 'media-pipeline-stalled'; new config blocks of the same names (null disables, omitted applies defaults).
  • New score reason keys servers may encounter: high-jitter, bandwidth-limitation, frozen-video, pixelated-video, audio-concealment, audio-time-stretch, high-jitter-buffer-delay, downscaled-screenshare. pixelated-video ranges 0–3.0, unlike every other normalized reason.
  • New doc docs/SCORE_CALCULATIONS.md — every reason key, threshold, ramp and formula, with what each key means for the user experience. The README’s scoring section is now the overview.

Source: GitHub release 4.7.0. Published at 2026-08-30T08:44:11Z (UTC).