Short answer
Cloud operations operates cloud resources with controlled changes, clear observability, and accountable cost. The expectations focus assessment on change artifacts, operating signals, recovery work, and cost decisions.
About Cloud operations
Operates cloud resources with controlled changes, clear observability, and accountable cost. The capability covers the ongoing operation of cloud environments rather than application feature delivery.
Use this competency for
- Roles that provision, change, monitor, recover, or optimize cloud resources.
- Functions accountable for cloud reliability, operational controls, resource ownership, or cost evidence.
Do not use this competency for
- Application roles with no responsibility for operating cloud resources or platform controls.
Important distinctions
Network administration
Network administration focuses on connectivity and traffic controls. Cloud operations covers the wider lifecycle of cloud compute, storage, platform services, observability, and cost.
Platform engineering
Platform engineering builds reusable developer capabilities. Cloud operations runs cloud resources and controls, even when no internal platform product is being built.
Expectations by level
IC1
Operates scoped resources
Completes well-defined cloud operations tasks with guidance, uses approved deployment and recovery steps, and verifies resource state after each change.
Observable behaviors
- Applies changes through the current controlled path.
- Checks required health and cost signals after a change.
- Records the owner, purpose, and result for changed resources.
Examples
- Updated a scheduled resource through the approved configuration and confirmed its next run.
- Removed an unused resource after verifying ownership and the recorded retention need.
IC2
Owns a cloud area
Independently operates a cloud service area, resolves ambiguous reliability or cost issues, and chooses rollout, monitoring, and recovery controls for changes.
Observable behaviors
- Plans cloud changes with rollback and validation criteria.
- Investigates service signals across resource dependencies.
- Uses usage and cost evidence to change capacity or resource design.
Examples
- Reduced an unexpected cost increase after tracing idle capacity to an outdated workload.
- Recovered a failed deployment using the tested fallback and documented the missing signal.
IC3
Sets cloud controls
Defines cloud operating standards across teams, frames systemic reliability and cost risks, and establishes patterns that teams can adopt with clear accountability.
Observable behaviors
- Defines shared controls for provisioning, observability, and recovery.
- Reviews cloud designs for failure paths, ownership, and cost exposure.
- Prioritizes cross-team improvements using incident, usage, and cost evidence.
Examples
- Introduced resource ownership and expiry rules after orphaned environments accumulated across teams.
- Set a shared recovery pattern for a managed service used by several products.