ClickHouse Cloud architecture
ClickHouse Cloud significantly simplifies operational overhead and reduces the costs of running ClickHouse at scale. There is no need to size your deployment upfront, set up replication for high availability, manually shard your data, scale up your servers when your workload increases, or scale them down when you’re not using them — we handle this for you. These benefits come as a result of architectural choices underlying ClickHouse Cloud:- Compute and storage are separated and thus can be automatically scaled along separate dimensions, so you don’t have to over-provision either storage or compute in static instance configurations.
- Tiered storage on top of object store and multi-level caching provides virtually limitless scaling and good price/performance ratio, so you don’t have to size your storage partition upfront and worry about high storage costs.
- High availability is on by default and replication is transparently managed, so you can focus on building your applications or analyzing your data.
- Automatic scaling for variable continuous workloads is on by default, so you don’t have to size your service upfront, scale up your servers when your workload increases, or manually scale down your servers when you have less activity
- Seamless hibernation for intermittent workloads is on by default. We automatically pause your compute resources after a period of inactivity and transparently start it again when a new query arrives, so you don’t have to pay for idle resources.
- Advanced scaling controls provide the ability to set an auto-scaling maximum for additional cost control or an auto-scaling minimum to reserve compute resources for applications with specialized performance requirements.
Capabilities
ClickHouse Cloud provides access to a curated set of capabilities in the open source distribution of ClickHouse. Tables below describe some features that are disabled in ClickHouse Cloud at this time.Database and table engines
ClickHouse Cloud is a highly-available, replicated service by default, built on the SharedMergeTree table engine family. When you create a table with a standard MergeTree-family engine, Cloud automatically substitutes the correspondingShared* engine. You don’t add a Shared or Replicated prefix yourself.
Replicated* engines are converted to the same Shared* equivalents. The substitution is visible in SHOW CREATE TABLE, which reports the Shared* engine even though your statement specified the plain variant.
The following table engines are also supported and used as written. The supported MySQL table engine is distinct from the unsupported MySQL database engine:
- URL
- View
- MaterializedView
- GenerateRandom
- Null
- Buffer
- Memory
IcebergS3andIcebergAzure- Deltalake
- Hudi
- MySQL
- MongoDB
- NATS
- RabbitMQ
- PostgreSQL
- S3
- Kafka
PaimonS3 and PaimonAzure table engines can be enabled on select ClickHouse Cloud services. Contact Support to confirm availability.
Interfaces
ClickHouse Cloud supports HTTPS, native interfaces, and the MySQL wire protocol. Support for more interfaces such as Postgres is coming soon.Dictionaries
Dictionaries are a popular way to speed up lookups in ClickHouse. ClickHouse Cloud currently supports dictionaries from PostgreSQL, MySQL, remote and local ClickHouse servers, Redis, MongoDB and HTTP sources.Federated queries
We support federated ClickHouse queries for cross-cluster communication in the cloud, and for communication with external self-managed ClickHouse clusters. ClickHouse Cloud currently supports federated queries using the following integration engines:IcebergS3andIcebergAzurePaimonS3andPaimonAzure(experimental; contact Support for availability)- Deltalake
- Hudi
- MySQL
- MongoDB
- NATS
- RabbitMQ
- PostgreSQL
- S3
User defined functions
SQL and executable user-defined functions are generally available in ClickHouse Cloud. See User-defined functions in Cloud.Settings behavior
This means:- Session-level settings (set via
SETstatement) aren’t propagated to UDF execution context - User profile settings aren’t inherited by UDFs
- Query-level settings don’t apply within UDF execution
Experimental features
Experimental features are disabled in ClickHouse Cloud services to ensure the stability of service deployments.Named collections
DDL-created named collections can be enabled on select ClickHouse Cloud services. Contact Support to confirm availability. Named collections defined in configuration files aren’t available because users can’t modify server configuration files in ClickHouse Cloud.Operational defaults and considerations
The following are default settings for ClickHouse Cloud services. In some cases, these settings are fixed to ensure the correct operation of the service, and in others, they can be adjusted.Operational limits
max_parts_in_total: 10,000
The default value of the max_parts_in_total setting for MergeTree tables has been lowered from 100,000 to 10,000. The reason for this change is that we observed that a large number of data parts is likely to cause a slow startup time of services in the cloud. A large number of parts usually indicate a choice of too granular partition key, which is typically done accidentally and should be avoided. The change of default will allow the detection of these cases earlier.
max_concurrent_queries: 1,000
Increased this per-server setting from the default of 100 to 1000 to allow for more concurrency.
This results in number of replicas * 1,000 concurrent queries for a service. A single-replica service supports up to 1000 concurrent queries regardless of tier. Multi-replica services on the Scale and Enterprise tiers support up to 1000 concurrent queries per replica.