RELIABILITY,
UPTIME AND COMMITMENTS.
We stand behind the stability and architectural endurance of the systems we engineer. This SLA details our incident classifications, response times, and engineering commitments.
Scope & Applicability
This Service Level Agreement (“SLA”) governs software systems, custom web applications, APIs, and cloud infrastructure maintained by EthosCore under an active Engineering Retainer or Managed Infrastructure Agreement.
For project-based engagements during active delivery sprints prior to production deployment, milestone acceptance criteria defined in individual Statements of Work apply.
Availability & Uptime Targets
EthosCore commits to delivering an architectural availability benchmark of 99.9% Uptime across production-grade microservices and database engines.
Maximum allowable monthly unscheduled downtime < 43 minutes.
Continuous automated ping and synthetic endpoint monitoring.
Automated on-call escalation via PagerDuty / Opsgenie.
Incident Severity Matrix
Incidents are categorized by operational business impact to ensure immediate priority for mission-critical disruptions:
Complete outage of the production application, core payment/POS billing failure, critical database corruption, or severe security breach preventing ongoing business transactions with no workaround.
Major business feature impaired (e.g. inventory sync lag, reporting generation failure), but the primary customer transaction flow remains operational through an identifiable alternative workflow.
Minor functionality defect, cosmetic UI rendering inconsistency, or administrative dashboard query timeout that does not prevent customer orders or data capture.
General technical questions, architectural advisory inquiries, scope refinement requests, or non-urgent configuration adjustments.
Response & Remediation Windows
| Severity | First Response SLA | Update Cadence | Target Mitigation |
|---|---|---|---|
| P1 (Critical) | < 60 Minutes (24/7) | Every 60 Minutes | Continuous until patched |
| P2 (Major) | < 4 Hours | Every 4 Hours | < 24 Hours |
| P3 (Moderate) | < 1 Business Day | Daily | Next Sprint Release |
| P4 (Request) | < 2 Business Days | As required | Backlog Prioritization |
Deployment & Rollback Protocol
Every production deployment executed by EthosCore engineers follows a zero-downtime, deterministic CI/CD release workflow:
- Automated Preview Environments: All pull requests are verified on isolated staging sandbox instances prior to merging into production branches.
- Instant Rollback Trigger: If automated health check endpoints or error-rate thresholds exceed 0.5% post-deployment, traffic is routed automatically to the previous immutable release image.
- Database Migration Safety: All schema migrations follow expand-and-contract patterns to ensure backward compatibility during transitions.
Scheduled Maintenance Windows
Routine infrastructure upgrades, major database version upgrades, and kernel security patches are conducted during pre-arranged maintenance windows:
Root Cause Analysis (RCA)
Following any P1 incident, EthosCore provides a comprehensive, blameless engineering post-mortem document within 5 business days of incident closure.
Dedicated Escalation Channels
Clients on engineering retainers have access to private direct channels for escalation: