- Learn
- ClickHouse
- Running ClickHouse on AWS
Running ClickHouse on AWS
Your options for running ClickHouse on AWS, from fully managed to self-hosted.
Overview
AWS is the most popular cloud platform for running ClickHouse. Whether you want a fully managed service or prefer to run your own infrastructure, AWS offers several paths.
This guide covers four main approaches: ClickHouse Cloud (the managed option from the ClickHouse team), self-managed on EC2, Kubernetes-based deployments on EKS, and AWS Marketplace offerings. Each comes with trade-offs in operational overhead, cost, and flexibility.
ClickHouse Cloud on AWS
ClickHouse Cloud is the managed service built and operated by ClickHouse, Inc. It runs on AWS infrastructure and handles provisioning, scaling, backups, and upgrades for you.
How it works
ClickHouse Cloud separates compute and storage. Your data lives in S3, and compute nodes scale independently. This means you can scale reads without duplicating data, and storage costs stay close to raw S3 pricing.
The service meters compute per minute in compute units of 8 GiB RAM and 2 vCPU. Storage is billed on the compressed size of data in your tables. You pay for what you use, with no need to pre-provision capacity.
Pricing tiers
ClickHouse Cloud offers three tiers:
Basic is designed for testing and starter projects. It provides a fixed single-replica service with 8 GiB RAM and 2 vCPU. Storage is capped at 1 TB of compressed data plus backups. There is no scaling (manual or automatic) and deployment is single-zone. Compute is about $0.22 per unit-hour in AWS us-east-1, so one 8 GiB replica running continuously costs roughly $159/month before storage, and less if the service idles.
Scale is the production workhorse. It supports multi-replica high availability, unlimited storage, and independent scaling of compute and storage. Automatic vertical scaling adjusts replica size based on load, and you can add replicas manually. Compute is about $0.30 per unit-hour in AWS us-east-1, so two 8 GiB replicas running continuously come to roughly $436/month before storage.
Enterprise targets mission-critical and compliance-heavy workloads. It adds custom hardware profiles, private regions, scheduled upgrades, customer-managed encryption keys, and HIPAA/PCI compliance features. Compute is about $0.39 per unit-hour in AWS us-east-1, so two 32 GiB replicas with 5 TB of storage come to roughly $2,400/month.
Storage is $25.30 per TB per month across all tiers in AWS us-east-1. Data egress (internet and cross-region) is billed separately. Rates change by cloud and region, so check the ClickHouse Cloud pricing page for current figures. New organizations get a 30-day trial with $300 in credits.
AWS regions
ClickHouse Cloud is available in multiple AWS regions, including us-east-1 (N. Virginia), us-west-2 (Oregon), eu-central-1 (Frankfurt), eu-west-1 (Ireland), and ap-southeast-1 (Singapore), among others. Pricing varies slightly by region, with us-east-1 as the baseline.
Setting up ClickHouse Cloud
- Sign up at clickhouse.com or through the AWS Marketplace (more on that below).
- Choose your tier and AWS region.
- Create a service. ClickHouse Cloud provisions your cluster and provides connection details.
- Connect using the native ClickHouse client, an HTTP endpoint, or a GUI tool like DB Pro.
The managed service is the fastest way to get started. It removes the need to manage infrastructure, plan capacity, or handle version upgrades.
Self-managed ClickHouse on EC2
Running ClickHouse on EC2 gives you full control over the database configuration, instance types, storage, and networking. This path makes sense when you need to customize ClickHouse settings that the managed service does not expose, or when your compliance requirements demand a self-hosted deployment.
Choosing instance types
ClickHouse is CPU-intensive for query processing and benefits from high memory for caching. The right instance family depends on your workload.
Memory-optimized (r-family): The r6i, r6g, and r7g families are a solid default. They provide a high memory-to-CPU ratio, which helps ClickHouse cache more data in memory and reduce disk reads. For cost efficiency, Graviton-based r7g instances offer strong performance at a lower price point than x86 equivalents. The newer r8g instances (Graviton4) offer up to 30% better performance than r7g, according to AWS.
Storage-optimized (i-family): The i3 and i4i families include local NVMe storage, which provides the lowest latency for disk-bound workloads. These are a good fit for customer-facing analytics where query latency needs to stay in the low milliseconds. The i4i family supports up to 30 TB of local NVMe per instance.
Compute-optimized (c-family): The c6i and c7g families work well if your workload is primarily CPU-bound (heavy aggregations, many concurrent queries) and your dataset fits largely in memory or on fast storage.
Here is a rough sizing guide:
| Workload size | Suggested instance | vCPU | RAM | Notes |
|---|---|---|---|---|
| Development | r6g.large | 2 | 16 GiB | Minimum for testing |
| Small prod | r7g.2xlarge | 8 | 64 GiB | Good for single-node setups |
| Medium prod | r7g.4xlarge | 16 | 128 GiB | Per node in a cluster |
| Large prod | i4i.8xlarge | 32 | 256 GiB | Local NVMe for low latency |
For clusters, plan on at least 3 nodes for high availability with ReplicatedMergeTree tables and a ClickHouse Keeper quorum.
Storage options
Storage choice has a big impact on both performance and cost.
EBS gp3 is the recommended default for most workloads. It provides a baseline of 3,000 IOPS and 125 MiB/s throughput, with the ability to provision up to 80,000 IOPS and 2,000 MiB/s per volume (IOPS scale at up to 500 per GiB, so the maximum needs a volume of at least 160 GiB). For instances with 32 vCPU or fewer, a single gp3 volume can saturate the available EBS bandwidth.
EBS io2 (and io2 Block Express) is the choice when you need guaranteed high IOPS. It supports up to 256,000 IOPS per volume and offers 99.999% durability. It costs significantly more than gp3, so reserve it for latency-sensitive workloads where gp3 throughput is not enough.
Instance store (NVMe) on i3 and i4i instances provides the lowest latency and highest throughput. The downside is that data on instance store is ephemeral. If the instance stops or terminates, the data is lost. You need to build replication (via ClickHouse's built-in replication) and plan for node replacement.
S3 for cold storage: ClickHouse supports tiered storage. You can configure a storage policy that keeps hot data on local or EBS volumes and moves older data to S3. This reduces storage costs for large datasets where older data is queried less frequently.
Installation steps
Here is a condensed setup for a single ClickHouse node on an EC2 instance running Ubuntu.
1. Launch an EC2 instance. Choose your instance type (e.g., r7g.2xlarge), attach a gp3 EBS volume (size depends on your data, 500 GiB is a reasonable start), and place it in your target VPC and subnet. Use an Amazon Linux 2023 or Ubuntu 22.04+ AMI.
2. Format and mount the EBS volume.
3. Install ClickHouse.
4. Configure ClickHouse. Edit /etc/clickhouse-server/config.xml to set the listen host, data paths, and logging. For a cluster, configure remote_servers and set up ClickHouse Keeper (or ZooKeeper) for coordination.
5. Start the service.
6. Verify the installation.
Cluster setup considerations
For production, run at least 3 ClickHouse nodes with ReplicatedMergeTree tables. You also need a ClickHouse Keeper ensemble (3 or 5 nodes) for coordination. Keeper can run on the same machines as ClickHouse for smaller deployments, or on separate instances for larger ones.
Place nodes across multiple Availability Zones for fault tolerance. Use a Network Load Balancer (NLB) or DNS round-robin to distribute client connections across nodes.
ClickHouse on EKS
Running ClickHouse on Amazon Elastic Kubernetes Service (EKS) is a good option if your organization already uses Kubernetes for other workloads. It provides a middle ground between fully managed and fully self-hosted: you manage the Kubernetes cluster, but an operator handles ClickHouse-specific concerns.
The ClickHouse Kubernetes operators
Two operators are commonly used:
Altinity Kubernetes Operator is the most mature option. It manages ClickHouse clusters through Custom Resource Definitions (CRDs). You describe your desired cluster state in a YAML file, and the operator handles provisioning, scaling, upgrades, and configuration changes. It applies changes using rolling procedures to minimize downtime.
Official ClickHouse Operator was released by ClickHouse, Inc. in 2025 under the Apache 2.0 licence. It manages both ClickHouse and ClickHouse Keeper clusters through its own custom resources. It is newer and still on 0.x releases, so check the documentation before relying on it in production.
Deployment approach
Altinity provides a Terraform module (terraform-aws-eks-clickhouse) that creates an EKS cluster optimized for ClickHouse with EBS storage and autoscaling. This is the fastest path to a production-ready setup.
The high-level steps:
- Provision an EKS cluster (Terraform, eksctl, or the AWS console).
- Install the ClickHouse operator via Helm.
- Define a
ClickHouseInstallationcustom resource with your desired cluster topology. - Apply the resource. The operator creates StatefulSets, Services, and PersistentVolumeClaims.
For storage on EKS, use the EBS CSI driver with gp3 volumes. Set the storageClassName in your PersistentVolumeClaim to point at a StorageClass backed by gp3.
EKS deployments are well suited for teams that want to standardize on Kubernetes tooling for all their stateful services. The trade-off is the added complexity of managing a Kubernetes cluster, node groups, and the operator lifecycle.
EKS node group recommendations
For ClickHouse pods, create a dedicated managed node group with the instance types discussed in the EC2 section (r7g or i4i families). Use taints and tolerations to ensure ClickHouse pods land on these nodes and are not preempted by other workloads.
Set resource requests and limits carefully. ClickHouse expects to have consistent access to memory and CPU. Do not oversubscribe ClickHouse nodes. A good rule of thumb is to set the ClickHouse pod's resource requests equal to its limits, guaranteeing a "Guaranteed" QoS class in Kubernetes.
For Keeper pods, smaller instances (m7g.large or t4g.xlarge) are sufficient. Keeper is lightweight but sensitive to latency, so avoid burstable instances for production Keeper nodes.
AWS Marketplace offerings
The AWS Marketplace provides several options for running ClickHouse, each with different levels of management.
ClickHouse Cloud (Pay-As-You-Go)
ClickHouse Cloud is available directly through the AWS Marketplace. Signing up this way lets you consolidate billing through your existing AWS account. You get the same service as signing up directly with ClickHouse, but charges appear on your AWS bill. This is useful for organizations that have committed AWS spend or Enterprise Discount Programs.
ClickHouse Cloud (Committed Contract)
For predictable workloads, you can commit to a specific spend amount through the Marketplace. Committed contracts often come with discounted rates compared to pay-as-you-go pricing.
Altinity.Cloud
Altinity offers a managed ClickHouse service through the AWS Marketplace. It runs in your own AWS account (a "bring your own cloud" model), giving you more control over data residency and networking. Altinity is a long-standing contributor to the ClickHouse ecosystem and maintains the Kubernetes operator.
AWS Quick Start (CloudFormation)
AWS published a Quick Start CloudFormation template for deploying a ClickHouse cluster on EC2, but its repository was archived in October 2024 and the guide still targets ClickHouse 23.3. Treat it as a reference for VPC and instance layout rather than a current install path, and pair it with a supported ClickHouse release.
Networking considerations
VPC design
Place your ClickHouse nodes in private subnets. They do not need public IP addresses. Clients should connect through a load balancer, VPN, or VPC peering.
For a self-managed cluster, all ClickHouse nodes and Keeper nodes must be able to communicate with each other on the required ports:
| Port | Protocol | Purpose |
|---|---|---|
| 8123 | HTTP | HTTP interface |
| 8443 | HTTPS | HTTP interface (TLS) |
| 9000 | TCP | Native protocol |
| 9440 | TCP | Native protocol (TLS) |
| 9009 | TCP | Inter-server replication |
| 9181 | TCP | ClickHouse Keeper |
Security groups
Create a security group for ClickHouse nodes that allows:
- Inbound on ports 8123/8443 and 9000/9440 from your application subnets or load balancer.
- Inbound on port 9009 from other ClickHouse nodes (for replication).
- Inbound on port 9181 from ClickHouse and Keeper nodes (for coordination).
- Outbound to S3 endpoints (for tiered storage or backups).
Keep rules tight. Do not open ClickHouse ports to the internet. If external access is needed, use AWS PrivateLink or a bastion host.
PrivateLink with ClickHouse Cloud
ClickHouse Cloud supports AWS PrivateLink, which lets you connect to your managed ClickHouse service over the AWS private network. This avoids sending traffic over the public internet, improving both security and latency. PrivateLink is available on the Scale and Enterprise tiers.
Load balancing
For self-managed clusters, place a Network Load Balancer (NLB) in front of your ClickHouse nodes. NLBs handle TCP traffic well and support the native ClickHouse protocol on port 9000. Configure health checks against the HTTP interface (/ping on port 8123 returns "Ok." when the node is healthy).
For HTTP-based access, an Application Load Balancer (ALB) works too. Set up a target group with health checks on the /ping endpoint.
Storage deep dive
EBS vs instance store
The choice between EBS and instance store depends on your durability and performance requirements.
EBS provides durable, network-attached storage. Data survives instance stops and restarts. Snapshots enable point-in-time backups to S3. The trade-off is that network-attached storage adds latency compared to local disks.
Instance store (local NVMe on i3, i4i, and similar families) provides the highest throughput and lowest latency. It is ideal for workloads where every millisecond counts. The risk is that data on instance store is lost if the instance stops, terminates, or experiences a hardware failure.
A common production pattern: use instance store for the primary data path and rely on ClickHouse replication (ReplicatedMergeTree with 2+ replicas) for durability. If a node fails, the data is rebuilt from its replicas. This gives you NVMe performance without sacrificing durability at the cluster level.
Backup strategies
For EBS-based deployments, take regular EBS snapshots. Automate this with Amazon Data Lifecycle Manager (DLM).
For any deployment, ClickHouse supports BACKUP and RESTORE commands that write to S3:
Schedule backups with a cron job or a Lambda function triggered on a schedule.
For ClickHouse Cloud, backups are handled automatically. The service takes daily backups and retains them based on your tier. You can also trigger on-demand backups through the ClickHouse Cloud console.
RAID configurations
For EC2 instances with multiple EBS volumes, you can stripe volumes into a RAID 0 array to increase throughput beyond what a single volume provides. This is useful on larger instances where a single gp3 volume's 2,000 MiB/s throughput cap becomes a bottleneck.
RAID 0 doubles throughput but offers no redundancy at the volume level. Rely on ClickHouse replication for data durability when using this configuration.
Monitoring with CloudWatch
Key metrics to track
If you are running ClickHouse on EC2 or EKS, use CloudWatch (along with the CloudWatch agent) to monitor infrastructure-level metrics:
- CPU utilization: ClickHouse queries are CPU-intensive. Sustained high CPU usage indicates you need more compute capacity or query optimization.
- Memory utilization: Track RSS memory usage. ClickHouse will use available memory for caching, but out-of-memory conditions cause query failures.
- Disk IOPS and throughput: Monitor EBS read/write IOPS and throughput to ensure you are not hitting volume limits.
- Network throughput: High network traffic can indicate excessive data shuffling between nodes or large result sets.
ClickHouse system tables
ClickHouse exposes internal metrics through system tables. These are more useful than CloudWatch for database-level monitoring:
Setting up CloudWatch dashboards
Install the CloudWatch agent on each ClickHouse node and configure it to collect system-level metrics (CPU, memory, disk I/O). For ClickHouse-specific metrics, use a Prometheus exporter (such as clickhouse_exporter or the built-in Prometheus endpoint in ClickHouse) and forward metrics to CloudWatch via the CloudWatch agent's StatsD or Prometheus scraping support.
Create alarms for:
- CPU utilization above 80% for 5 minutes.
- Available disk space below 20%.
- Replication lag above a threshold (query
system.replicasforabsolute_delay). - Failed queries per minute (query
system.query_logfor errors).
Alternative monitoring stacks
Many teams pair ClickHouse with Grafana for dashboards. ClickHouse has a native Grafana data source plugin, and you can query the system tables directly to build operational dashboards. This approach often provides more detail than CloudWatch alone.
For a more integrated setup, consider running the ClickHouse Prometheus endpoint (enabled in config.xml under <prometheus>) alongside the CloudWatch agent. This gives you both infrastructure metrics in CloudWatch and detailed ClickHouse metrics in Grafana, all without third-party exporters.
Cost estimates
These are rough monthly estimates for running ClickHouse on AWS in us-east-1, based on list prices as of September 2026 and 730 hours per month. Actual costs depend on reserved instance pricing, data transfer, and workload patterns.
ClickHouse Cloud
| Tier | Configuration | Estimated monthly cost |
|---|---|---|
| Basic | 1 x 8 GiB replica, 24/7 | ~$159 |
| Scale | 2 x 8 GiB replicas, 24/7 | ~$436 |
| Enterprise | 2 x 32 GiB replicas, 5 TB | ~$2,406 |
Storage adds $25.30/TB per month across all tiers (included in the Enterprise row above). Data egress is billed separately.
Self-managed on EC2 (on-demand pricing)
Small deployment (single node, development or light production):
| Component | Spec | Monthly cost |
|---|---|---|
| EC2 instance | r7g.2xlarge | ~$313 |
| EBS gp3 | 500 GiB | ~$40 |
| Total | ~$353 |
Medium deployment (3-node cluster with replication):
| Component | Spec | Monthly cost |
|---|---|---|
| EC2 instances (3x) | r7g.4xlarge | ~$1,876 |
| EBS gp3 (3x) | 1 TiB each | ~$246 |
| Keeper instances (3x) | t4g.medium | ~$74 |
| Total | ~$2,196 |
Large deployment (6-node cluster, NVMe storage):
| Component | Spec | Monthly cost |
|---|---|---|
| EC2 instances (6x) | i4i.4xlarge | ~$6,014 |
| Keeper instances (3x) | m7g.large | ~$179 |
| S3 cold storage | 10 TiB | ~$236 |
| Total | ~$6,428 |
Reserved instances (1-year, no upfront) or Savings Plans can reduce EC2 costs by 30-40%. Graviton-based instances are typically 10-20% cheaper than their x86 counterparts for comparable performance.
Cost comparison
For small to medium workloads, ClickHouse Cloud is often the better value when you factor in the engineering time needed to manage a self-hosted deployment. Patching, monitoring, capacity planning, and backup management add up quickly.
Self-managed deployments become more cost-effective at larger scales, where the per-node cost savings outweigh the operational overhead. The crossover point depends on your team's experience with ClickHouse and infrastructure management.
Data transfer costs
Data transfer is one of the most overlooked costs on AWS. Keep these in mind when planning your ClickHouse deployment:
- Same-AZ traffic between EC2 instances is free. If your ClickHouse nodes and application servers are in the same AZ, replication and query traffic incur no transfer charges.
- Cross-AZ traffic costs $0.01/GB in each direction. A 3-node cluster spread across 3 AZs will generate cross-AZ traffic for replication. For write-heavy workloads (e.g., ingesting 1 TB/day with 2 replicas), this can add $600+/month.
- Internet egress starts at $0.09/GB. If ClickHouse serves data to clients outside AWS, this adds up fast. Use VPC endpoints for S3 access to avoid NAT gateway egress charges on tiered storage reads and backups.
- ClickHouse Cloud egress is billed by ClickHouse, not AWS. In AWS us-east-1 it starts at about $0.115/GB for public internet egress and $0.031/GB for inter-region transfer. See the pricing page for other regions.
To minimize transfer costs, co-locate your application and ClickHouse in the same region and, where possible, the same Availability Zone. For multi-AZ clusters, accept the cross-AZ cost as the price of high availability.
IAM and security
IAM roles for EC2
Assign an IAM instance profile to your ClickHouse EC2 instances. This avoids hardcoding AWS credentials in ClickHouse configuration files. The instance role should grant:
s3:GetObject,s3:PutObject, ands3:ListBucketfor the S3 buckets used by tiered storage and backups.kms:Decryptif you use KMS-encrypted S3 buckets.cloudwatch:PutMetricDataif the CloudWatch agent publishes custom metrics.
Encryption
Enable encryption at rest for all EBS volumes. AWS encrypts volumes using AES-256 with keys managed by KMS. This is transparent to ClickHouse and adds negligible latency.
For data in transit, configure ClickHouse to use TLS on all client-facing ports (8443 for HTTPS, 9440 for the native protocol). Between cluster nodes, enable TLS for the interserver replication port (9009) as well. Store TLS certificates in AWS Secrets Manager or use AWS Certificate Manager for rotation.
ClickHouse Cloud handles encryption at rest and in transit by default, with no additional configuration required.
Best practices summary
Start with ClickHouse Cloud if you want to move fast and do not have specific requirements that demand self-hosting. It is the lowest-effort path to production.
Choose EC2 when you need full control over configuration, want to use instance store for maximum performance, or have compliance needs that require running in your own account with specific networking controls.
Choose EKS when your team already operates Kubernetes and wants to standardize tooling. The Altinity operator is mature and handles most operational tasks.
Storage: Use gp3 as your default EBS choice. Provision IOPS and throughput as needed. Consider instance store (i4i) for latency-critical workloads, backed by replication for durability.
Networking: Keep ClickHouse in private subnets. Use PrivateLink for ClickHouse Cloud. Use security groups to restrict port access to known sources.
Monitoring: Combine CloudWatch for infrastructure metrics with ClickHouse system tables for database-level visibility. Set up alerts for CPU, disk, and replication lag.
Backups: Use EBS snapshots for volume-level backups. Use ClickHouse's BACKUP command for logical backups to S3. Test your restore process regularly.