Server-Guard · Intelligent Ops Engine

24/7 Guard After Launch

Around-the-clock monitoring and fast response stop risks before users feel them.

Watch Pipeline: From Alert to Closure

  1. 01

    Inspect

    31 checks on schedule, collected in parallel — anomalies surface immediately.

  2. 02

    Alert

    Prometheus / Zabbix wired in; dedupe-and-enrich staging keeps alert storms out.

  3. 03

    Diagnose

    LLM root-cause analysis plus playbook rules — structured, actionable fixes.

  4. 04

    Heal

    Low-risk fixes run automatically over SSH — rollback-ready, verified after repair.

  5. 05

    Approve

    Humans sign off on high-risk moves via mobile or console — AI never oversteps.

server-guard console
Server-Guard watch console — live system
Server-Guard watch console — live system

31-Item Automated Inspection Matrix

3131 checks · three categories

  • System resources

    • CPU health (dual sampling)
    • System load
    • Memory
    • Swap
    • Disk mount points
    • Process / port liveness
    • Kernel OOM / panic
    • systemd services
  • Application layer

    • JVM heap & GC
    • Nginx config & error logs
    • Redis slow queries
    • Docker containers & image age
    • MySQL status
    • Java threads
  • Security audit

    • OS lifecycle (EOL early warning)
    • Login-capable user inventory
    • Zombie processes

Scheduled with distributed locks for parallel collection; parsers adapt — present components get checked, absent ones skipped, no false alarms.

Four-Level Risk Response

  • LOW

    Low

    Handling
    Auto-approved
    Typical Scenario
    Log cleanup, temp-file cleanup, non-critical service restarts
  • MEDIUM

    Medium

    Handling
    Human approval (mobile notice)
    Typical Scenario
    Database restarts, config changes with reload
  • HIGH

    High

    Handling
    Human approval (dual-channel confirm)
    Typical Scenario
    Multi-service impact, data changes, batch operations
  • CRITICAL

    Critical

    Handling
    Console approval + mandatory rollback plan
    Typical Scenario
    Operations that could interrupt service or lose data

Three-layer risk scoring with time awareness — risk levels auto-escalate during peak hours.

Human-in-the-Loop Approval

Humans make the call on critical actions: a four-state machine keeps everything on record — approve anytime from mobile IM or the web console.

  1. Created
  2. Notified
  3. Pending
  4. Approved / Rejected
  • Auto-reject on timeout
  • Crash-recoverable
  • Dual-channel notice & approval

Defense-in-Depth Security

  • Hard command blocking

    7 danger classes, 21+ fatal commands permanently blocked — filesystem damage and privilege opens never run.

  • 5-tier roles, deny by default

    Scoped from read-only to batch execution with regex whitelists; anything not explicitly allowed is denied.

  • Full-chain audit trail

    Every inspection, alert, diagnosis, approval and execution is persisted — traceable up to 365 days.

AI Diagnostic Core

  • Dual-LLM diagnosis

    Dual providers with automatic failover and circuit breaking; batch diagnosis with independent root-cause analysis.

  • Playbook rules

    Built-in playbooks for disk, CPU, Nginx and more — rule matching and stepped execution alongside the LLM.

  • Multi-agent orchestration

    Five agent roles — inspect, alert, diagnose, execute, batch — scheduled per scenario with decomposition, concurrency and retries.

  • RAG knowledge retrieval

    Semantic retrieval over past cases and ops manuals — every decision informed by history.

Zero external dependencies

  • No MySQL / Redis / message queues
  • SQLite + JSONL out of the box
  • Single-node deployment, minimal ops cost
  • One node guards dozens of servers

Certificate Lifecycle Management

  1. 01

    Tiered expiry warnings

    Warnings at 30/14/7 days and on expiry, plus drift and missing-cert detection on a schedule.

  2. 02

    ACME auto-renewal

    A self-built RFC 8555 client renews automatically — private CAs and manual uploads as fallback.

  3. 03

    Gated deployment

    High-risk gates with atomic replacement, three-stage checks and automatic rollback on failure.

  4. 04

    Private-key safety

    0600/0700 permission enforcement (fail-closed), 7 rolling backups, fully audited access.

Real Disposal Scenarios

  • Disk alert, handled at 3 a.m.

    Monitoring pushes a disk-at-95% alert; AI blames log bloat, risk scores low, cleanup runs automatically — disk recovers with zero human touch.

  • Java OOM, one-tap diagnosis

    An inspection catches old-gen at 95%, queues a heap-dump request for approval — tap "approve" on your phone and it runs. No SSH needed.

  • Bad config, caught in time

    AI suggests a reload, but syntax checks fail: the safety layer flags it, blocks execution and pushes approval with a rollback plan — an outage averted.

  • Certificates, renewed on schedule

    Warning 30 days out, ACME renews automatically; after approval, atomic replacement and health checks run — with rollback if anything fails.

Core Capabilities

  • Around-the-clock monitoring

    Unified monitoring across systems, apps and business metrics — anomalies reach the right person immediately.

  • Fast incident response

    Tiered response and playbooks in place — recovery first, retrospective after.

  • Routine inspection and maintenance

    Patches, backups, certificates and capacity checked on schedule — risks removed early.

  • Security hardening

    Baseline checks and security patches kept current to shrink the attack surface.

  • Capacity and cost tuning

    Continuous observation and tuning keep performance steady and costs lean.

Delivery Process

  1. 1

    Assessment

    Take stock of assets and risk baselines; define the care scope and priorities.

  2. 2

    Monitoring onboarding

    Stand up monitoring and alerting with full visibility of key metrics.

  3. 3

    Playbook setup

    Set tiered response and change management, with duties and timings agreed in advance.

  4. 4

    Daily care

    Inspections, watch duty and tuning run as a routine; anomalies are handled by the book.

  5. 5

    Review and improvement

    Regular health reports feed a rolling loop of retrospectives and improvements.

Get a Proposal