Skip to content

Monitoring

Monitoring gives integrators the evidence needed to support users: health state, recent errors, operation logs, and job progress.

Rokks separates immediate service health from user workflow outcomes. A healthy service can still return an error for one malformed request. A failed job can affect one source while the rest of the platform remains healthy.

Operational visibility

  1. Start from the reported user action.
  2. Identify the environment and tenant.
  3. Check job records for the affected time.
  4. Check service health for the route family involved.
  5. Review recent errors and correlate them with the job or request.
  6. Resolve the dependency, data, or configuration issue.

Do not declare the platform healthy solely because the home page loads.

Do not ignore job errors that users have stopped watching. Background failures still need operational follow-up.