DEV Community

Cover image for Sentinel: The Watchdog and Derived Points
Philip Shaw
Philip Shaw

Posted on Originally published at glitchedpixel.io

Sentinel: The Watchdog and Derived Points

A door sensor that only reports when the door opens or closes can go months without sending anything. If nothing has happened, that silence is correct. If the sensor has died, the silence looks identical, and no freshness window can tell the two apart. So the registry won't accept an on_change point unless it also declares a heartbeat or explicitly sets never. Whoever adds the point has to say how anyone would know it had died, and the answer is recorded in a commit.

A derived point can mislead in the same way. Computed naively, a 15-minute mean of a flow temperature keeps producing numbers after its sensor goes quiet, and a mean built from a single reading looks entirely reasonable on a dashboard. So the output carries a quality worked out from its inputs. A window with fewer than min_samples readings comes out bad, with the reason insufficient_samples. A stale input makes the output stale unless the point has opted into live_or_stale. And a derived point that stops emitting misses its own deadline and is marked stale by the same watchdog that watches the devices.

Restarts are where silence is easiest to misread. Each point's deadline is recomputed from its last observed_at when the core starts, so nothing has to be persisted. Without a startup grace, though, a fleet that reports every five minutes would go stale within seconds of the restart. A core running behind on ingest is worse: every point would pass its deadline while its data sat queued on the bus, and fail-safe interlocks would assert across the whole estate. Staleness is paused in both cases, and the lag pause is announced on its own signal so it can't hide a real fault. The slow part of a planned restart is rebuilding derived-point windows, and running those queries on a small pool, with one query shared by every derived point that reads the same input, keeps it bounded.

Staleness and derived values both reach the rules engine as ordinary transitions on ordinary points, so the engine, two parts from now, only has to understand one kind of input.

What this document owns: freshness policies and the single timer wheel that enforces them, the heartbeat requirement on on_change points, liveness as a freshness policy, how staleness fires, how it recovers and how it ranks against unavailable, startup grace, suppression during ingest lag and the three limits on it, flap damping, the derived point descriptor and its operator set, quality propagation, evaluation order and the cap on chain depth, windows and the retention guard that bounds them, the cold-start rehydration budget, emit policy, and the restricted expression grammar with its two binding contexts. It does not own the lag mechanism, its arithmetic or the scale targets, which belong to Core Runtime; the retention classes, their policy values and how each rolls up, which belong to Retention & Compaction; the raw-table query a window rehydrates from, which is defined in Ingest Path & Log Schema; or how a rule triggers on a transition, claims an output and uses the grammar in its when clause, which belong to Rules Engine.

These two land together because the watchdog's output is just another observation, and a derived point with no fresh inputs is exactly the staleness problem again one level up.

Watchdog

A single timer wheel in the core, not a timer per point. Every point with a freshness policy has one pending deadline; when an observation lands, the deadline is cancelled and rescheduled. A hierarchical wheel at 1 s resolution handles tens of thousands of points on one thread and stays O(1) on reschedule, which is the operation that runs constantly. At the four thousand points targeted in Core Runtime this is not close to a constraint.

Deadlines survive restart because they are derived, not stored: observed_at + window + grace, recomputed for every point during projection load. Nothing to persist.

Policies, from the descriptor:

  • periodic(interval, grace) — deadline at observed_at + interval + grace. Grace defaults to half the interval, capped at 60 s.
  • on_change(max_silence) — deadline at observed_at + max_silence. Requires the device to heartbeat.
  • ttl(duration) — hard expiry, no grace.
  • never — no deadline. Config values, static setpoints.

on_change without a heartbeat is the trap. A door sensor that reports only on transition is silent for months when nothing happens, and there is no window that distinguishes "quiet" from "dead". The registry rejects on_change unless the point also declares heartbeat: <interval> or explicitly sets never. Forcing that choice at registration is the whole point: the question "how would I know if this died?" gets answered once, in a commit, rather than never.

For BLE advertisement points, periodic against the advertising interval is usually more honest than on_change, since advertisements repeat regardless of change.

Liveness is a freshness policy, not a value. A point that reports "still running" every second is not a measurement; the information is in the gap, not in the sample. It is modelled as an ordinary point with a tight periodic policy, with a rule triggering on quality rather than on value. The system already treats absence as a first-class signal, so nothing extra is needed — and a rule watching for true to become false would be watching the wrong thing, since a dead process does not report false.

Firing writes a quality_changed entry to control_event and publishes a transition on state.{point_id}. Same envelope shape as a value transition, changed: false, value untouched. To the rules engine it is indistinguishable in structure from any other transition, which is the design goal.

stale and unavailable interact by precedence: an availability event from a driver sets unavailable and cancels the point's deadline. When availability returns, the point stays unavailable until real data arrives — no synthetic promotion. This stops a flapping device from generating alternating stale/unavailable noise.

Recovery is not automatic. A point goes stale → live only when an observation arrives that satisfies its policy. An observation with observed_at older than the current deadline — a buffered replay landing late — does not clear staleness, though it is still logged. This is the second place observed_at versus ingested_at earns its keep.

When staleness is suppressed

Two conditions pause staleness evaluation entirely. Both exist because the absence of data is sometimes a fact about the core rather than about the device, and reporting it as device staleness is a lie that causes the system to act.

Startup grace. For the first N seconds after core start, restored → stale transitions are suppressed, where N is the point's window plus grace. Otherwise every point in a fleet that reports every 5 minutes goes stale within seconds of a restart and floods the log. The suppression is per point, so a genuinely dead device still surfaces at the right time.

Ingest lag. While the core is behind, all staleness evaluation pauses. Freshness is computed from observed_at, so a core ten minutes behind would take every point in the system past its deadline while the data sits queued and healthy on the bus — and rules with on_input_loss: release would drop their claims, and fail-safe interlocks would assert, across the whole estate, because ingest was slow. That is close to the worst available response and it is entirely self-inflicted.

Three constraints on the lag suppression, because withholding a signal must not become a way to hide problems:

  • The threshold is computed, not configured. Suppression engages when lag exceeds the smallest window + grace in the registry. Past that point staleness is unreliable for at least one point, and there is no principled reason to trust it for the others.
  • It is global, not per point. Unlike startup grace, the condition is a property of the core, so all staleness evaluation pauses rather than pausing selectively.
  • The condition is loud. Lag is a first-class signal, sustained lag is alertable, and core.{n}.staleness-suppressed says plainly that this is happening. The system is not going quiet; it is replacing a false signal with a true one.

When lag clears, deadlines are recomputed from each point's observed_at and genuinely dead devices surface within one window. The lag mechanism and its arithmetic live in Core Runtime.

Flap damping. A point crossing live/stale repeatedly is marked flapping rather than emitting every transition. Transitions are counted in a rolling window; past a threshold, a flapping control event is emitted and further transitions coalesce until it settles. This is a rules-engine concern as much as a watchdog one — a rule triggering on staleness of a flapping point will fire endlessly.

Derived points

A derived point is a point. Same ID scheme, same descriptor fields, same quality states, same freshness, same bus subject. Consumers cannot tell the difference and should not be able to.

id            boiler.ch1.flow-temp-15m
kind          measurement
quantity      temperature
unit_native   Cel
subject       water
subject_ref   zone:upstairs-ch
derivation:
  op          mean
  inputs      [ boiler.ch1.flow-temp ]
  window      15m
  emit        on_input | interval(1m)
  min_samples 3
  input_quality  live_only | live_or_stale
  freshness   periodic(1m, grace=30s)
Enter fullscreen mode Exit fullscreen mode

The descriptor declares quantity and unit explicitly rather than inheriting them, because some operators change both. A rate over energy yields power in W; a delta over temperature yields temperature_delta, which is a different quantity precisely because it is not affine-convertible. If the engine cannot verify that the declared output quantity matches what the operator produces from the input quantity, that is a registry validation failure at load time.

Operator set, deliberately small:

op notes
mean, min, max, sum, count mean rejected when the input quantity is not agg_safe
median, p95 for noisy analogue sources
stddev a staleness-adjacent signal — a stuck sensor has zero variance
delta last minus first in window; output quantity differs from input
rate delta over time; output quantity differs
debounce(duration) boolean stable for a duration
hysteresis(on, off) boolean from a numeric input
expression restricted arithmetic over named inputs; grammar below

hysteresis as a derived point is the most consequential of these. Thresholds belong here, not in rules. A rule that says "pump on when flow-temp-low is true" is readable and testable; a rule carrying its own threshold pair is neither, and the same threshold used by three rules is three chances to get it inconsistent.

Quality propagation is where most implementations of this go wrong. The rule:

  • No input has ever been live → output unknown
  • Any required input unavailable → output unavailable
  • Any required input bad → output bad
  • Window has fewer than min_samples live samples → output bad, reason insufficient_samples
  • Any required input stale and policy is live_only → output stale
  • Otherwise → live

The min_samples case matters more than it looks. A 15-minute mean computed from one sample is not a 15-minute mean, and it will look completely reasonable on a dashboard. min_samples defaults to a figure derived from window and expected input interval — around 50% of expected — rather than to 1.

input_quality: live_or_stale exists for windows long enough that a missed poll should not invalidate them. It is opt-in, and the output carries lower confidence, but it stops a 24-hour average collapsing because one reading was late.

Evaluation. Derived points form a DAG. Acyclicity is validated at registry load and cycles rejected there rather than detected at runtime. Evaluation is in topological order within a single transition batch, so a chain of three derived points settles in one pass rather than three bus round trips.

Chain depth is capped at three or four levels. Deeper than that is a rules problem wearing a derived-points costume.

Windows are queries, not buffers. As established in the ingest design, rehydration is a SELECT against the raw table for the input's retention class. The in-memory window is a cache of that query, and it can be rebuilt at any time. This is what makes derived points restart-transparent, and it is worth protecting: no derived-point state that cannot be recomputed from the log.

Windows are bounded by both duration and sample count, so a misconfigured 5 Hz sensor with a 24-hour window does not consume the core.

Windows are bounded by the retention of their inputs

A window longer than the retention behind it rehydrates from a truncated range: it produces a value computed over less time than it claims, which min_samples will usually but not always catch.

Retention differs by class, so the check is per input, against that input's class, not against a single global figure. A derived point reading a short-retention diagnostic must satisfy the diagnostic horizon even where it would comfortably pass a measurement one. That is a strictly better check than the global version and it catches the case most likely to be wrong.

Two validations, and the second is the one that will actually bite:

  • At registry load, reject any derived point whose window exceeds its input's class retention less one chunk interval, and warn above half of it. Zero slack is not enough, because a chunk dropping partway through a rehydration truncates the window silently. The check reads the retention policy from the database rather than a constant in the code — retention is a deployment parameter, and hardcoding a figure here means changing it later breaks derived points with no warning at all.
  • At retention change, refuse a shortening that would orphan an existing window. Retention gets shortened when a disk fills, which is exactly the moment nobody is thinking about derived points. This is the same class of guard as refusing to remove a vocabulary member that points still reference.

A window that cannot fit takes the rollup as its input instead. A thirty-day mean cannot rehydrate from raw samples dropped at fourteen days, but it can read the hourly aggregate, which is kept far longer and is exactly the right granularity for a window that long. The choice is explicit in the descriptor and enforced at load, rather than discovered at the next restart when the window comes back short.

Rehydration is the cold-start budget

At the scale targets in Core Runtime — a thousand or more derived points — this step is essentially all of the planned-restart time, and it is the only part of cold start that grows with configuration rather than with fleet size. Three things keep it bounded:

  • Bounded concurrency, not sequential. A thousand round trips executed one at a time is minutes of latency spent waiting rather than working. Run on a small pool, the wall time collapses toward the slowest few queries.
  • Share queries across derived points reading the same input. With a thousand derived points over two thousand inputs, several derived points will read the same source over overlapping windows. Group by input, issue one query covering the widest window in the group, and slice in memory. This is where the largest saving is, and it costs nothing at registry load to compute the grouping.
  • Watch the compression boundary. A window reaching past the compression age decompresses chunks to serve, which is far slower than a hot-chunk read. Windows that cross it are worth knowing about; a long window is not wrong, but it should be a deliberate choice rather than a surprise at the next restart.

Rehydration time is measured as a self point from the first restart. It is the number that degrades quietly as the rule set grows, and the one that makes an outage worse than it needs to be, since recovery is cold start plus drain.

Emit policy. on_input recomputes and publishes on every input transition, which is right for debounce and hysteresis where latency matters. interval(period) recomputes on a schedule, which is right for aggregates over slow-moving inputs where per-sample recomputation is wasted work. Aggregates generally use interval, and the interval becomes the basis for the derived point's own freshness policy.

Derived points respect their own change_policy deadband exactly as device points do, so a moving average that moves by 0.001 °C does not emit a transition.

The restricted expression grammar

Defined here and referenced by the rules engine. There is one grammar with two binding contexts, not two grammars, and it is written down before either consumer is built because an informal grammar is how eval gets in.

Legal:

  • arithmetic + - * / % and unary minus
  • comparison == != < <= > >=
  • boolean and or not
  • min, max, abs, clamp, round
  • numeric, boolean, quality-member and enum-member literals
  • names bound by the enclosing declaration, carrying exactly two accessors: .value and .quality

Rejected at load, not at runtime: function definitions, lambdas, loops, comprehensions, indexing, general attribute access, imports, string formatting, any call not in the list above, and any name not declared in the enclosing inputs.

.value and .quality are an enumerated accessor pair, not general attribute access. The distinction has to be explicit, because "no attribute access" and flow.quality == live otherwise contradict each other. The accessor set is closed, it has two members, and adding a third is a grammar change rather than a convenience.

Quality members are legal only as operands of == and !=. Ordering comparison on quality is meaningless — stale > live is not a question — and the parser rejects it rather than coercing something.

Binding contexts. In a derived expression point, names bind to input values and a bare x means x.value; .quality is available but rarely wanted, because quality propagation is handled by the rules above rather than inside the expression. In a rule's when, names bind to the full (value, quality) pair and there is no bare-name sugar: a rule that reads a value without considering its quality is usually a bug, and it should have to say so in writing.

Both contexts parse to an AST at registry load, using one parser. Two parsers is how the contexts diverge.

The property that keeps the rules engine small

Derived points feed derived points, and staleness of a derived point is computed by the same watchdog with the same policies. A derived point whose inputs went quiet stops emitting, misses its own deadline, and goes stale — and a rule can trigger on that without knowing or caring that the point was derived.

That uniformity is what lets the rules engine stay small. Rising edge, falling edge, moving average, staleness, fresh data and bad data are each either a transition predicate or a derived point, and the engine only needs to understand transitions over points.

Exit criteria

No spurious edges at cold start. Load a checkpointed projection, start the core, and assert zero restored → stale transitions inside each point's grace, and the correct ones after it. This is the whole reason startup grace exists, and it is easy to get subtly wrong in a way that only shows up as a burst of log entries nobody reads.

Staleness suppression engages and clears. Simulate lag past the computed threshold and assert that no staleness transitions are emitted, that core.{n}.staleness-suppressed is true, and that when lag clears the deadlines are recomputed from each point's observed_at so a genuinely dead device surfaces within one window. The failure mode to catch is suppression that never lifts.

Quality propagation, one fixture per row of the table. Including the two that produce plausible wrong answers rather than errors: a window below min_samples yields bad with insufficient_samples rather than a mean of two readings, and a non-agg_safe input rejects mean at load rather than computing one.

The window guard fails the build. A derived point whose window exceeds its input's class retention is rejected at registry load, and separately, a retention shortening that would orphan an existing window is refused. Both directions, because the second is the one that happens under pressure.

Rehydration is measured, and the grouping demonstrably works. Record the wall time, and assert that the number of queries issued is materially below the number of derived points — if it equals it, the shared-query grouping is not doing anything and cold start will degrade linearly as the rule set grows.


On Monday - the Dev Diary. The latest on where the specification and the build disagree as the system takes shape, and which of the two had to change.

Then the Wednesday after - The Command Plane. Where Sentinel starts writing to devices: desired and reported held as separate values, a durable claim set with banded arbitration, and a state machine that follows a fire-and-forget write through to a verified outcome.

Start of the series: An Introduction. The map, the two decisions every later document is downstream of, and why a specification at this scale is being published in public while the system it describes gets built.

Top comments (0)