Class RuntimeStatsSnapshot
A point-in-time read of the numbers an operator watches: tick percentiles, overruns, the durability wait, per-system cost, entities per archetype, and the session figures when replication is running. Produced by ReadStats(int).
public sealed class RuntimeStatsSnapshot
- Inheritance
-
RuntimeStatsSnapshot
- Inherited Members
Remarks
Why this exists beside the STATS wire block. The same values were already computed once a second and sent to subscribing game clients, and
to nothing else: the encoder holds them in a private array and its only exit is a frame. So a curl, a CLI verb, a dashboard or a log line could
not read one of them. Worse, the computation is gated twice over — it happens only inside a subscriptions runtime, and only when the APPLICATION's
catalog declares metrics — so an engine with replication off, or one whose app declared none, computed nothing at all. This reads the sources directly
and therefore answers on any running engine.
Read at human rate, not on the tick path. It allocates (one snapshot, a few arrays) and walks the telemetry window, so it belongs behind an HTTP endpoint, a CLI verb or a log timer — not in a system. Every figure outside Realms is either an instantaneous level or a percentile over the window the caller asked for, so it needs no differencing.
Realms is the exception, and deliberately. Its work figures are cumulative since the process started, because the question they exist to answer is "has this realm EVER cost anything" — a realm nobody is in is supposed to contribute exactly zero, and a per-tick level cannot tell "zero this tick" from "zero always". Difference two reads for a rate.
The telemetry ring is a single-writer diagnostic structure with no publication protocol, so a sample read while the tick driver is writing it may
be torn. The consequence is one perturbed sample in a percentile, which is why the ring is read here rather than mirrored into something that would have
to be written on the tick path to gain nothing. The STATS encoder accepts the same trade for the same reason.
Constructors
RuntimeStatsSnapshot()
public RuntimeStatsSnapshot()
Properties
Archetypes
Live entity count per registered archetype, at the moment of the read.
public ArchetypeStat[] Archetypes { get; init; }
Property Value
DurabilityWaitP99Ms
99th-percentile per-tick durability wait over the window, in milliseconds: the Unit-of-Work flush, which in WAL mode is RequestFlush followed
by WaitForDurable.
public double DurabilityWaitP99Ms { get; init; }
Property Value
Remarks
Read beside TickP99Ms, this is the "compute versus durability" split an operator needs: a tick at 12 ms of which 9 is this one is a disk problem, and the same 12 ms with 0.2 here is a compute problem. It is a per-TICK figure, not per-commit.
NetOutBytesTotal
Bytes the send pump has written since the process started. Cumulative: difference two reads to get a rate.
public long NetOutBytesTotal { get; init; }
Property Value
Overruns
Ticks in the window whose duration exceeded the target. A count, not a rate — read it against TicksInWindow.
public int Overruns { get; init; }
Property Value
Remarks
Against TargetTickMs, the 1× base-rate target — not against the multiplier-adjusted budget a throttled tick is actually given. That
is the same display semantic the scheduler's OverrunRatio carries, and deliberately so (the #289 follow-up in
DagScheduler.Telemetry.cs): the effective ratio exists to drive the overload detector, because a workload that fits comfortably at multiplier N
still exceeds the 1× target and a detector reading it that way never deescalates. An operator's question is the other one — "is the engine holding the
rate it was configured for" — and the answer to that is measured against the configured rate.
The consequence to know: while TickMultiplier is above 1, ticks counted here include ones doing exactly what the overload manager told them to. A rising count with a rising multiplier is load being shed, not the engine falling behind; a rising count at multiplier 1 is the engine falling behind. The two readings need each other, which is why the multiplier is on this snapshot at all.
RealmPassSteps
Realms visited by a serial per-realm replication stage since the engine opened: each stage adds the number of realms served when it ran. Cumulative — difference two reads.
public long RealmPassSteps { get; init; }
Property Value
Remarks
Realms D-7 asks whether the serial share of a tick grows with the realms being SERVED or with the realms merely registered. This counts the answer instead of inferring it: hold the served set at one realm, register two thousand more, and this number must not move. Nothing derivable from Realms can make that check — a stage that walked the registry would serve the same set while doing two thousand times the work.
The parallel per-chunk stages are not in it (the index merge's and far fold's chunk dispatch, and the projection's per-chunk sort), because a counter on a path several workers run at once measures the counter. Their cost is read from the clock, in the D-7 bench.
Divided by the ticks in the same interval, a steady state reads (stages + 1/SweepEvery) × served — the unplaced sweep is the one stage that
runs one tick in SweepEvery rather than every tick. How many stages is deliberately not promised here: two of them are skipped on a tick
with no bound push session, and one runs once per pushed archetype, so a caller that needs the constant measures it at one served count and holds later
counts to it rather than writing a number down.
Zero means replication is off, not that the engine has one realm. The hub serves realm 0 from construction, so a single-realm engine that replicates anything still counts it. That is the opposite of Realms, which IS empty for a single-realm engine.
RealmPolicyEvaluations
Realm-policy evaluations since the engine opened — one per tick that ran, each deciding every registered realm's run state and divisor. Cumulative — difference two reads.
public long RealmPolicyEvaluations { get; init; }
Property Value
Remarks
The companion to RealmPassSteps and the honest half of it: policy is O(registered realms) by design, not by accident, because a
dormant realm is precisely one whose state has to be re-decided in case a session arrived. What must hold is that it runs once per tick: a
second evaluation per dispatch, or per realm, would multiply the one term here that is allowed to scale with registration.
Realms
One row per registered realm, in registration order: what it is, what it is doing, and what its sessions have cost.
public RealmStat[] Realms { get; init; }
Property Value
Remarks
Per registered realm, not per served one, and that distinction is the point. A realm nobody is in has no replication state at all — the hub drops it — so its counters do not read zero, they do not exist. Reporting only the served realms would make "this realm cost nothing" indistinguishable from "this realm is not in the list", which is exactly the claim a realm world lives or dies on: an SWG galaxy registers a realm per enterable building and a thousand of them are empty at any moment. Served says which case a zero row is.
Empty when the engine has one realm, where every figure above already describes it.
ReplicatedArchetypes
Archetypes the application declared for replication. Zero is how a reader knows the figures below are zero because nothing is replicated, rather than because a replicating server happens to be idle.
public int ReplicatedArchetypes { get; init; }
Property Value
Remarks
A count rather than a bool, and deliberately not "is the subscriptions runtime alive": a subscriptions runtime is built on every
Start() whether or not an application declared anything, so its existence answers "has this runtime started", which the
caller already knows. What it wants to know is whether there is anything to replicate.
ReplicationPrologueMsTotal
Milliseconds spent in the frame stage's single-threaded prologue since the engine opened, and the ticks it was measured over. Cumulative — difference two readings. Both zero unless phase timing is enabled, which it is not by default.
public double ReplicationPrologueMsTotal { get; init; }
Property Value
Remarks
The companion to ReplicationTrackP99Ms and a different quantity: the track's percentile covers the whole stage, most of which is chunks on worker threads, while this is the part no second core can help with. The difference between N sessions in one realm and the same N sessions in N realms, divided by N, is the per-realm fixed cost of everything inside this window — which is what Realms D-7 was measured with.
It does not contain every serial per-realm step, and a D-7-style reading has to say so. The frame prologue holds the overload step, the index build and world order, the far flushes, the LOD census and the unplaced sweep. Three more serial per-realm stages sit outside it, in other tracks' prepare phases: the index merge's plan, the index finish and far plan, and the queued shadow checks. So this is a lower bound on the serial per-realm cost, not the whole of it; the upper bound is the whole tick.
Cumulative on purpose, and the tick count comes with it. The engine also keeps this as a mean over the whole run, which is the wrong shape for a measurement: a run whose configuration changed half way through reports one number for both halves. The ticks here are the ticks the frame stage was busy, not every tick the engine ran, so they are the right denominator and they are not derivable from Tick.
ReplicationPrologueTicks
Ticks the frame stage was busy, the denominator for ReplicationPrologueMsTotal. Zero unless phase timing is enabled.
public long ReplicationPrologueTicks { get; init; }
Property Value
ReplicationTrackP99Ms
99th-percentile cost of the replication track itself over the window, in milliseconds.
public double ReplicationTrackP99Ms { get; init; }
Property Value
Sessions
Sessions currently open — CCU.
public int Sessions { get; init; }
Property Value
Systems
Mean duration of each system over the window, in microseconds, in schedule order. Empty when the ring is empty.
public SystemStat[] Systems { get; init; }
Property Value
TargetTickMs
public double TargetTickMs { get; init; }
Property Value
Tick
The tick the window ends at — the newest tick the ring had recorded when the snapshot was taken.
public long Tick { get; init; }
Property Value
TickMultiplier
The tick-rate multiplier the newest tick in the window ran under: 1 at the configured rate, 2 or more while the overload manager is modulating.
public int TickMultiplier { get; init; }
Property Value
Remarks
Present so Overruns can be read correctly — see its remarks. The newest tick's rather than the window's maximum, because it answers "what is the engine doing now"; a window that changed multiplier mid-way is visible as a mismatch between this and the overrun count, which is a truer signal than either a maximum or a mean would give.
TickP50Ms
Median tick duration over the window, in milliseconds.
public double TickP50Ms { get; init; }
Property Value
TickP99Ms
99th-percentile tick duration over the window, in milliseconds.
public double TickP99Ms { get; init; }
Property Value
TicksInWindow
How many recorded ticks the percentiles below were computed over. Zero means the ring was empty and every percentile is zero.
public int TicksInWindow { get; init; }