Skip to content

feat(analytics): add W64 runtime history and SLO APIs - #245

Open
ammarheidari wants to merge 1 commit into
mainfrom
w64-operational-analytics-runtime
Open

ammarheidari wants to merge 1 commit into
mainfrom
w64-operational-analytics-runtime

Conversation

@ammarheidari

@ammarheidari ammarheidari commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Authority

W64 #213 is ACTIVE. W63 historical metrics is complete and protected-main verified.

Backend completion slice

  • live consumer lag evidence from the governed Kafka consumer read-view;
  • consume-rate remains explicit Unavailable when no real metrics provider exists;
  • bounded consumer-lag sampler writes real lag observations into W63 history when historical metrics are enabled;
  • W63-backed consumer history adapter replaces the unavailable history port only when history is configured;
  • bounded historical consumer-lag trend query;
  • SLO evaluation over complete available trend points only;
  • live/trend/SLO HTTP APIs protected by existing ConsumerRead authorization;
  • unavailable/unknown evidence never fabricates numeric zero;
  • no broker/topic throughput is invented when no provider exists.

Safety

  • sampler is bounded by cluster/group/concurrency/interval policy;
  • missing or failed observations remain missing;
  • no mutation, offset commit, hidden consumer-group ownership or raw payload access;
  • history queries retain W63 query ceilings;
  • partial/unknown evidence is propagated explicitly.

Exact head: e55f28ee985238ac0aac6f4f2b53f69752ceb74d.

Refs #213.

Canonical runtime reconciliation

  • protected-main parent: fd9dc052cfc0e31b6b56d06a7147f462ce2f709a;
  • exact head: a6a47f53214eae49ca7a5c4ffb82c4eb693d0158;
  • one DCO-signed commit;
  • bounded consumer-lag sampling writes real Kafka lag evidence into W63 only when history is configured;
  • historical diagnostics/trends read from W63 with provider ceilings;
  • live lag precision loss becomes Partial rather than falsely exact;
  • stale consume-rate evidence becomes Stale;
  • Unknown historical evidence cannot make a trend Available;
  • SLO uses only complete Available points and reports Partial/Unavailable/Unknown truth explicitly;
  • broker/topic throughput remains unavailable rather than fabricated without a real provider;
  • fresh exact-head CI, CODEOWNER approval and Codex review required.

Final exact-head replay

  • protected-main parent: fd9dc052cfc0e31b6b56d06a7147f462ce2f709a;
  • exact head: e55f28ee985238ac0aac6f4f2b53f69752ceb74d;
  • one DCO-signed commit;
  • all runtime/history/SLO truth and precision fixes are included;
  • source is frozen pending fresh exact-head CI/review.

Copy link
Copy Markdown
Contributor Author

@codex review

Please review exact W64 backend head c07eca273e591de7d53752ae458090327160205c, especially truthful unavailable metrics, bounded lag sampling, W63 history mapping and SLO evidence semantics.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-28T19:57:22.050712Z e55f28e Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copy link
Copy Markdown
Contributor Author

@codex review

Please review current W64 runtime source head 7d9d208559ae11db46e08b7564370f35f61826d7 after the test compile correction. Focus on bounded sampling, truthful unavailable evidence, history mapping and SLO semantics.

@ammarheidari
ammarheidari force-pushed the w64-operational-analytics-runtime branch from 99d6528 to a6a47f5 Compare September 28, 2026 19:49

Copy link
Copy Markdown
Contributor Author

@codex review

Please review canonical exact W64 backend head a6a47f53214eae49ca7a5c4ffb82c4eb693d0158, especially bounded lag sampling, historical truth mapping, stale/precision semantics and SLO evidence calculation.

akhiabanchian
akhiabanchian previously approved these changes Sep 28, 2026
Refs #213

Signed-off-by: Ammar Heidari <ammar@arad-itc.org>
@ammarheidari
ammarheidari force-pushed the w64-operational-analytics-runtime branch from 9ca46f9 to e55f28e Compare September 28, 2026 19:53

Copy link
Copy Markdown
Contributor Author

@codex review

Review final exact W64 backend head e55f28ee985238ac0aac6f4f2b53f69752ceb74d only. Source is frozen; please focus on bounded sampler behavior, history provider truth and SLO semantics.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e55f28ee98

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +223 to +224
builder.Services.AddHostedService<
ConsumerLagHistorySamplingHostedService>();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Coordinate lag sampling across HA replicas

When historical metrics run in HighAvailability mode, every API replica registers this hosted service, so each replica independently issues up to 500 lag reads per cluster and writes another set of samples every interval. The per-process semaphore therefore does not provide a deployment-wide bound, and the slightly different observation timestamps also make these duplicate samples distinct rows that can bias later SLO calculations. Gate sampling with a shared lease/fence, as the historical maintenance worker does, or otherwise nominate a single sampler.

Useful? React with 👍 / 👎.

Comment on lines +469 to +476
var state =
points.Length == 0
? OperationalTrendState.Unknown
: points.Any(point =>
point.State is not
OperationalEvidenceState.Available)
? OperationalTrendState.Partial
: OperationalTrendState.Available;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Mark truncated history as partial before evaluating SLOs

When a trend contains more than MaxPoints, the store returns raw.Truncated = true, but this state calculation can still label the trend Available if every returned point is available. EvaluateSloAsync then computes compliance from only that truncated prefix and also reports the SLO as available, while OperationalSloResult carries no truncation field to reveal the omitted evidence. Include truncation in the partial-state decision so bounded queries cannot produce apparently complete SLO results.

Useful? React with 👍 / 👎.

Comment on lines +561 to +567
var state =
usable.Length !=
trend.Points.Count ||
trend.State ==
OperationalTrendState.Partial
? OperationalEvidenceState.Partial
: OperationalEvidenceState.Available;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Treat gaps in the requested SLO window as partial evidence

If sampling succeeds only once during a requested hour—for example because Kafka reads failed for the other intervals—the returned history contains one available point and this logic reports the SLO as fully Available, because it compares usable points only with points that happened to be returned. No check establishes coverage of the requested fromUtc/toUtc window or detects missing sampler intervals, so a sparse success can yield misleading 0% or 100% compliance. Use the coverage timestamps and expected sampling cadence to mark incomplete windows partial or unknown.

Useful? React with 👍 / 👎.

Comment on lines +116 to +124
var selected =
groups.Value
.OrderBy(group =>
group.GroupId,
StringComparer.Ordinal)
.Take(
_policy
.MaxGroupsPerCluster)
.ToArray();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Rotate the capped consumer-group sample

For clusters with more than MaxGroupsPerCluster groups, sorting and taking the same prefix on every cycle permanently excludes every group after the cap. Those groups can still be queried through the new trend/SLO endpoints but will never acquire history, even across arbitrarily many sampling intervals. Keep the per-cycle bound while rotating or cursoring the selected subset so all groups eventually receive samples.

Useful? React with 👍 / 👎.

Comment on lines +163 to +165
var exact =
total <=
LargestExactlyRepresentableInteger;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Compare lag precision against an integer bound

When total is 9_007_199_254_740_993, comparison with this double constant first converts the long to double, rounding it down to 9_007_199_254_740_992; exact consequently becomes true. The same conversion rounds the stored gauge value, so the sampler records an inexact value as Stable, unlike the live path which correctly uses a long bound. Make the precision limit a long to preserve the intended comparison.

Useful? React with 👍 / 👎.

Comment on lines +442 to +456
var raw =
await _history.QueryAsync(
new HistoricalMetricQuery(
metricName,
query.Resource.ClusterId,
OperationalMetricHistoryNames
.ResourceKind(
query.Resource.Kind),
query.Resource.ResourceId,
query.FromUtc,
query.ToUtc,
MaxSeries: 1,
query.MaxPoints),
cancellationToken)
.ConfigureAwait(false);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Return unavailable evidence when the history provider fails

When a configured history store times out or becomes unavailable, this call propagates the exception; the trend and SLO endpoints catch only ArgumentException, so clients receive a generic 500 instead of the explicit unavailable evidence returned when no provider is configured. Catch provider timeout/unavailability failures here, while preserving caller cancellation, and return an Unavailable trend so configured-but-failing providers obey the same evidence-truth contract.

Useful? React with 👍 / 👎.

Comment on lines +590 to +599
var state =
point.State switch
{
"Partial" =>
OperationalEvidenceState.Partial,
"Unknown" =>
OperationalEvidenceState.Unknown,
_ =>
OperationalEvidenceState.Available,
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve non-available historical point states

Any provider state other than the exact strings Partial and Unknown falls through to Available. A valid stored point marked Stale (or another non-stable provider state) is therefore presented as available and is included in SLO compliance, changing provider truth rather than conservatively preserving it. Map Stable explicitly to available and map recognized stale/unavailable states appropriately, treating unknown state strings as partial or unknown rather than available.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants