Build a monitoring agent that detects and diagnoses four test incidents
Project brief
Build a runnable monitoring/diagnostics workflow for one RenX test service or agent workflow. Monitor an agreed health endpoint and sanitized event/log feed, detect problems, and generate actionable alerts with source evidence. This is one bounded operational pilot, not a replacement for the production monitoring system. Propose the signals, runtime and checks in your application.
We will provide an isolated test target and define four incident scenarios before work: unavailable endpoint, sustained excessive latency, a burst of errors and a failed/stalled agent job. Detect all four within two minutes of the agreed trigger. Each alert must identify the affected component, timestamps, supporting events and a useful next diagnostic step. Where evidence is insufficient, mark the cause unknown rather than inventing one.
Include deduplication and a recovery notification. A four-hour healthy replay must produce no incident alerts. Demonstrate one safe recovery in the sandbox, such as restarting a disposable fixture, only after an explicit approval step. Do not restart production services, change permissions, delete records or expose secrets. Use least-privilege read-only monitoring access; any sandbox repair is separately allowlisted.
Deliver code/configuration, setup instructions and incident/recovery evidence so RenX can run the workflow again. Acceptance is based on measured detection and safe behavior, not on a written monitoring plan. Production rollout and continuous on-call coverage are excluded.
Delivery deadline: 2026-09-30. Planned work window: up to 12 calendar days after contracting and receipt of the agreed inputs, within this deadline. Confirm a feasible start date, exact scope and required materials before contract acceptance. If a later start needs a different deadline, agree and record the revised deadline before accepting; no work is requested before agreement. RenX supplies the agreed authorized or synthetic materials before work begins. Additional tools, model usage and other third-party costs require separate agreement.
Deliverables & acceptance
What you'll deliver
- Runnable monitoring/diagnostics workflow for one approved test target.
- Alert templates, deduplication/recovery logic and editable configuration.
- Four incident demonstrations, healthy replay and approval-gated sandbox recovery evidence.
What the result must meet
- All four agreed injected incidents are detected within two minutes, measured against recorded trigger times.
- Alerts identify component, time and supporting evidence, suggest a next check and do not assert unsupported root causes.
- No incident alerts occur in the agreed four-hour healthy replay; repeated signals for one incident are deduplicated and recovery is reported.
- One allowlisted sandbox recovery succeeds only after approval; unauthorized recovery attempts are blocked and no production changes occur.