The ROSA HyperFleet is a strategic redesign of Red Hat OpenShift Service on AWS (ROSA) with Hosted Control Planes (HCP). This project transforms ROSA from a globally-centralized management model to a regionally-distributed architecture where each AWS region operates independently with its own control plane infrastructure.
Key Goals:
- Regional Independence: Each region operates autonomously with its own cluster lifecycle management service to reduce global dependencies
- Operational Simplicity: GitOps-driven deployment with zero-operator access model
- Modern Cloud-Native Architecture: Built on AWS services (EKS, RDS, API Gateway)
- Disaster Recovery: Declarative state management with cross-region backups
-
Regional Cluster (RC) - EKS-based cluster running core services:
- Platform API (customer-facing with AWS IAM auth)
- Hyperfleet Operator - postgres-based cluster lifecycle controller
- kube-applier - DynamoDB-backed resource distribution to MCs
- ArgoCD - GitOps deployment
- Tekton - infrastructure provisioning pipelines
-
Management Clusters (MC) - EKS clusters hosting customer control planes:
- Run HyperShift operators hosting multiple customer control planes
- Dynamically provisioned and scaled per region
- Private Kubernetes APIs with no network path to RC (ideal state)
-
Customer Hosted Clusters - ROSA HCP clusters with control planes in MC
- Compute: Amazon EKS (Regional + Management Clusters)
- Networking: VPC, API Gateway (regional), VPC Link v2, ALBs
- Storage: Amazon RDS (hyperfleet-db), Amazon ElastiCache Valkey (rate limiting), EBS volumes
- Identity: AWS IAM for authentication and authorization
- Infrastructure: Terraform modules with GitOps patterns
- CI/CD: ArgoCD (apps), Tekton (infrastructure pipelines)
- Resource Distribution: kube-applier (DynamoDB-backed controller applying resources to MCs)
- Languages: Go (primary backend), Shell scripting
- Container Orchestration: Kubernetes via EKS
Work for the ROSA HyperFleet is tracked in Jira under two parent Outcomes:
- HPSTRAT-62 ("Red Hat Cloud Data Sovereignty"): Feature-driven work covering the regional platform build-out (architecture, infrastructure, services, tooling).
- HPSTRAT-11 ("FedRAMP Moderate Technical Delivery"): Compliance work covering FedRAMP security controls, audit requirements, and certification readiness.
Day-to-day engineering tasks (epics, stories, bugs) live in the ROSAENG project under the [ROSA] HyperFleet team (customfield_10001, id 0c538cd9-152b-49f6-ad7c-e2fa2f865809).
When creating ROSAENG issues:
- Set the team via
additional_fields:{"customfield_10001": "0c538cd9-152b-49f6-ad7c-e2fa2f865809"} - Do not set a component
- ALWAYS use the architect agent for changes to:
docs/design/- Any architectural decisions or patterns
- Use adversary agent for security review of code changes (supply chain, infrastructure, application security)
- Use code-reviewer agent for code quality review
- Use ci-troubleshooter agent for diagnosing CI/CD failures
- Use documentation-updater agent for reviewing documentation freshness
- GitOps First: ArgoCD for cluster configuration management, infrastructure via Terraform
- Private-by-Default: EKS clusters use fully private architecture with ECS bootstrap
- Declarative State: Hyperfleet Operator (backed by hyperfleet-db) maintains single source of truth for all cluster state
- Event-Driven: kube-applier reads desire documents from DynamoDB (written by hyperfleet-operator) and applies them to MCs via hyperfleet-dynamo GSI polling
- Regional Isolation: Each region operates independently with minimal cross-region dependencies
- Explicit Feature Flags: Optional or environment-specific infrastructure (e.g., CloudTrail, PagerDuty, resources with per-account limits) should be gated behind
enable_*configuration flags. Avoid patterns like checking against the environment's name to change behavior or functionality.- Feature flags should default to what keeps the best developer experience — focus on the lowest barrier to getting a new region started. We'd rather have verbose production configs than require developers to understand every flag just to get going.
- Bootstrap Strategy: Use ECS Fargate for private EKS cluster bootstrap (see
docs/design/fully-private-eks-bootstrap.md) - No Public APIs: All EKS clusters are fully private with VPC-only access
- ArgoCD Self-Management: Clusters manage their own ArgoCD installations via GitOps
- Rate Limiting: Per-account rate limiting at the Platform API layer using GCRA algorithm with ElastiCache Valkey as shared counter store (see
docs/design/rate-limiting-architecture.md)
terraform/
├── modules/eks-cluster/ # EKS with private bootstrap
├── modules/ecs-bootstrap/ # Fargate bootstrap tasks
└── config/ # Cluster configuration templates
argocd/
├── config/ # Live Helm chart configurations
│ ├── management-cluster/ # MC application templates
│ ├── regional-cluster/ # RC application templates
│ └── shared/ # Shared configurations
└── README.md
.chai-bot/ # Chai Bot scheduled tasks
├── rosa_hyperfleet_adversary_scan.md # On-demand single-repo adversary security scan
├── rosa_hyperfleet_ci_daily_health_report.md # Daily CI health report
├── rosa_hyperfleet_ci_weekly_status.md # Weekly Jira epic progress + PR stats
├── rosa_hyperfleet_docs_update.md # Weekly doc staleness detection & update PRs
└── rosa_hyperfleet_weekly_security_report.md # Weekly consolidated adversary scan across all repos
docs/
├── README.md # Architecture overview
├── FAQ.md # Architecture decisions Q&A
├── design/ # ADRs (Architecture Decision Records)
├── environment-provisioning.md
├── hostedcluster-provisioning.md # User-facing: create & access a hosted cluster
├── hostedcluster-teardown.md # Admin-only: manual teardown & force cleanup
├── development-environment.md
├── adding-component-pre-merge.md
└── sop/ # Standard operating procedures
- Update Terraform modules in
terraform/modules/ - Run
make pre-pushbefore committing or pushing — this is the all-in-one command that runsterraform-fmt,check-docs,check-rendered-files,helm-lint, andterraform-validate, matching the full CI suite. Individual targets (e.g.make terraform-fmt) can still be used for targeted runs. - Ensure architect agent reviews any architectural changes
- Update ArgoCD configurations in
argocd/ - Follow GitOps patterns - ArgoCD will sync changes
- Test in development region first
- Run
make pre-pushbefore pushing —check-docs(prettier markdown) and other non-Terraform checks apply to all change types
- Add region config to
config/<environment>/and render withuv run scripts/render.py - Bootstrap the central pipeline (see
docs/environment-provisioning.md) - ArgoCD bootstrap handles core service deployment
- Management Clusters auto-provision as needed
- Run
make pre-pushbefore pushing to validate all rendered files and documentation
IMPORTANT: Files in deploy/ are generated from config/ templates. Never edit them directly.
When you modify any file in config/<environment>/ (e.g., config/stage/defaults.yaml, config/integration/us-east-1.yaml):
-
Automatically regenerate rendered files by running:
uv run scripts/render.py
This regenerates all files in
deploy/<environment>/, including:_merged_config.yaml— merged view of the config hierarchyargocd-values-*.yaml— ArgoCD Helm valuesargocd-bootstrap-*/applicationset.yaml— ApplicationSet templatespipeline-*-inputs/terraform.json— Terraform input variables
-
Verify changes with
make check-rendered-filesormake pre-push -
Commit both the config source files and the regenerated deploy files together
The render script merges config from the inheritance chain:
flowchart TD
A[config/defaults.yaml] --> B["config/<env>/defaults.yaml"]
B --> C["config/<env>/<region>.yaml"]
C --> D["deploy/<env>/<region>/*"]
Never commit changes to deploy/ files without regenerating them from config/ sources.
See docs/development-environment.md for full usage — provisioning, resync, E2E, teardown, and port forwarding.
- AWS IAM Only: Use AWS IAM for all authentication/authorization
- Private Networking: No public endpoints except regional API Gateway
- Least Privilege: Follow AWS IAM best practices for service roles
- Encryption at Rest: KMS-encrypted EKS secrets, RDS, and EBS volumes
- Network Segmentation: Dedicated security groups for VPC endpoints and services
- High Availability: Multi-AZ NAT Gateways eliminate single points of failure
- Break-Glass Access: Use ephemeral containers for emergency access only
- Markdown: All markdown files must be formatted with
prettier. Runnpx prettier --write '**/*.md'before committing markdown changes. - Diagrams: Always use Mermaid for diagrams in markdown files, never ASCII art.
- Pre-push (required): Run
make pre-pushbefore committing or pushing. This single command runs the full CI validation suite —terraform-fmt,check-docs(prettier markdown),check-rendered-files,helm-lint, andterraform-validate— matching exactly what CI will enforce. Skipping this step is how formatting and lint failures reach CI (e.g. PR #364). - Terraform Validation: Always run
terraform validateandterraform plan - Format Check:
make terraform-fmt(also run automatically bymake pre-push) - ArgoCD Health: Verify applications sync successfully
- Security Review: Use architect agent for security-sensitive changes
Platform alerting and recording rules are defined as PrometheusRule CRs in the alerting-rules chart (argocd/config/regional-cluster/alerting-rules/templates/). Rules are evaluated by Thanos Ruler against Thanos Query. See docs/adding-alerting-rules.md for a developer guide on adding new rules, including the error budget burn rate pattern used for SLA alerts.
The Platform API implements per-account rate limiting using the GCRA (Generic Cell Rate Algorithm) via go-redis/redis_rate. Rate limit counters are stored in ElastiCache Valkey 9.1, chosen over Redis OSS for 20% lower cost, open-source licensing (BSD 3-Clause), and performance improvements.
Architecture: API Gateway (global throttle) → Platform API middleware (per-account GCRA) → ElastiCache Valkey (shared counters). Rate limiting runs after identity extraction but before authorization, so over-limit requests are rejected cheaply. The system fails open — if Valkey is unavailable, requests pass through.
Key files (rosa-hyperfleet):
| File | Purpose |
|---|---|
terraform/modules/elasticache-valkey/main.tf |
ElastiCache Valkey replication group, KMS, security groups, parameter group |
terraform/modules/elasticache-valkey/variables.tf |
cluster_id, vpc_id, node_type, engine_version |
terraform/modules/elasticache-valkey/outputs.tf |
Valkey endpoint and port outputs |
scripts/bootstrap-argocd.sh |
Threads Valkey endpoint from Terraform outputs to ECS bootstrap env vars |
terraform/modules/ecs-bootstrap/main.tf |
Writes redis_endpoint annotation on the ArgoCD cluster secret |
config/templates/argocd-bootstrap/applicationset.yaml.j2 |
Reads redis_endpoint annotation into Helm valuesObject |
argocd/config/regional-cluster/platform-api/values.yaml |
Rate limit config (routes, rates, burst), Valkey endpoint |
argocd/config/regional-cluster/platform-api/templates/ |
Deployment, ratelimit-configmap, servicemonitor templates |
argocd/config/regional-cluster/alerting-rules/templates/ |
PrometheusRule CRs for rate limit alerts |
docs/design/rate-limiting-architecture.md |
ADR: rate limiting design decisions |
Key files (rosa-hyperfleet-api): Rate limiting Go implementation lives in the API repo — pkg/ratelimit/ (GCRA limiter, config loading), pkg/middleware/ratelimit.go (HTTP middleware), cmd/rosa-regional-platform-api/main.go (wiring).
Data flow for Valkey endpoint: Terraform output → bootstrap-argocd.sh (combines host:port) → ECS task env var → cluster secret annotation → ApplicationSet valuesObject → Helm values → REDIS_ENDPOINT env var on platform-api pods.
Scheduled CI/documentation tasks run via Chai Bot. Schedules are defined in .chai-bot/.
Slack threading convention: Chai Bot uses delimiters to split output into a top-level channel message and threaded replies:
---THREAD_DETAILS---— everything after this line becomes threaded replies (not posted to the channel)---THREAD_BREAK---— separates individual threaded replies
Daily CI report tracks nightly-ephemeral and nightly-integration jobs. Top-level message shows today's status + 10-day trend table. Threaded replies are created only when jobs are failing, using the ci-troubleshooter agent for root cause analysis.
Weekly status groups epics by status (In Progress with completion %, To Do, Done) and summarizes PR activity across rosa-hyperfleet, rosa-hyperfleet-api, and rosa-hyperfleet-cli.
Docs update detects stale documentation across the three repos and creates PRs to fix it, using the documentation-updater agent for validation.
Adversary scan runs a weekly Groundwork-mode security scan of each repo using the adversary agent and posts severity-ranked findings to Slack. This is static/adversarial analysis, not CVE scanning.
Makefile- Standardized provisioning commandsbootstrap-argocd.sh- ECS Fargate bootstrap scriptargocd/config/shared/argocd/- ArgoCD self-management Helm chart- Design decisions follow ADR format in
docs/design/