enterprise-agent-ops

Helps developers run long-lasting AI agents safely with security checks and monitoring.

Installation
Run `npx skills add "https://github.com/affaan-m/ecc" --skill "enterprise-agent-ops"` to install this skill, then follow its SKILL.md instructions for my next request.

Paste this into Claude Code, Cursor, or any agent that can run commands.

What this skill does
What it does: - Manages AI agents that run for long periods in the cloud or on servers - Watches what the agents are doing with logs and measurements - Controls who can do what and sets safety limits - Keeps track of changes and can undo them if something goes wrong - Tracks how well the agents are working and how much they cost. When to use it: - You have an AI agent running all the time on a server - You need to watch what the agent is doing and make sure it stays safe - You want to update the agent without breaking things - You need to know if something goes wrong and fix it quickly - You want to measure how well the agent is working
SKILL.mdShow the author's original SKILL.md
---
name: enterprise-agent-ops
description: Operate long-lived agent workloads with observability, security boundaries, and lifecycle management. Use when running long-lived agent workloads that need observability, security boundaries, or lifecycle control.
metadata:
  origin: ECC
---

# Enterprise Agent Ops

Use this skill for cloud-hosted or continuously running agent systems that need operational controls beyond single CLI sessions.

## Operational Domains

1. runtime lifecycle (start, pause, stop, restart)
2. observability (logs, metrics, traces)
3. safety controls (scopes, permissions, kill switches)
4. change management (rollout, rollback, audit)

## Baseline Controls

- immutable deployment artifacts
- least-privilege credentials
- environment-level secret injection
- hard timeout and retry budgets
- audit log for high-risk actions

## Metrics to Track

- success rate
- mean retries per task
- time to recovery
- cost per successful task
- failure class distribution

## Incident Pattern

When failure spikes:
1. freeze new rollout
2. capture representative traces
3. isolate failing route
4. patch with smallest safe change
5. run regression + security checks
6. resume gradually

## Deployment Integrations

This skill pairs with:
- PM2 workflows
- systemd services
- container orchestrators
- CI/CD gates

Mirrored from the author's public source. Install counts from the open skills registry.

The systems behind these skills get built for partners every week.

Partner with us