> ## Documentation Index
> Fetch the complete documentation index at: https://private-7c7dfe99-parallel-read-in-order-multi-part.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

> Opt-in compiled Rust codec for ClickHouse Connect

# Rust native codec

ClickHouse Connect can decode query results and encode inserts with a compiled Rust codec instead of the default Python and Cython implementation. The Rust codec applies only to client-managed `FORMAT Native` traffic, which covers `query`, `query_np`, `query_df`, their block and row streaming variants, and inserts including `insert_df`. The Arrow methods use `FORMAT Arrow` and are unaffected, as are raw queries, raw inserts, and non-Native formats.

The codec is experimental and opt in. The Python codec remains the default.

<h2 id="installation">
  Installation
</h2>

The compiled codec ships as a separate wheel named `clickhouse-connect-core`, which provides the `_ch_core` extension module. Install the codec with the integrations your application uses:

```bash theme={null}
pip install "clickhouse-connect[rust]"         # Python rows and Native inserts
pip install "clickhouse-connect[rust,numpy]"   # NumPy results
pip install "clickhouse-connect[rust,pandas]"  # Pandas results and inserts
```

Core wheels are published for CPython 3.10 through 3.14 and free-threaded CPython 3.14t on Linux, macOS, and Windows. On free-threaded Python, importing the core extension doesn't re-enable the GIL.

Ordinary Rust NumPy and Pandas output doesn't require PyArrow or any other Arrow package. This includes `query_np`, `query_df`, their streaming variants, and `query(..., use_numpy=True)`. NumPy queries don't require Pandas. Add the `async` extra for the async client.

PyArrow is still required for Arrow methods and Arrow-backed Pandas storage. Install the `arrow` extra for those features:

```bash theme={null}
pip install "clickhouse-connect[rust,pandas,arrow]"
```

Extended Pandas string output follows Pandas' `mode.string_storage` option. When PyArrow is installed, strings use the existing Arrow conversion, with full UTF-8 validation, for either storage backend. Without PyArrow, the driver uses the Rust object conversion and Pandas' default storage uses Python strings. Explicitly requesting Arrow string storage without PyArrow raises a dependency error. Invalid UTF-8 keeps the existing hex rendering with either storage backend.

The codec, its packaging, and its dependency set are experimental. Direct buffer conversion now handles non-nullable 8-64-bit integers, Float32/64, Boolean, BFloat16, Interval, Date, Date32, DateTime, and DateTime64 columns. Scalar Time and Time64 columns, including nullable values, also use direct buffers, as do nullable BFloat16 and extended Pandas nullable integer, float, and Interval output. BFloat16 retains float32 output and its existing Pandas Float32 missing-value policy. Interval values remain signed counts in their declared unit. Temporal conversion retains declared units and existing timezone rules. Nullable SimpleAggregateFunction aliases of Time64 now return correct durations and NaT for NULL values. Nullable scalar DateTime64(9) columns also use buffers to preserve nanoseconds in NumPy and Pandas output. Other nullable dates and timestamps retain their object conversion paths.

Arrays of Time/Time64, including nested arrays and nullable leaves, and LowCardinality(Time), including nullable values, also use direct buffers. Their nesting, duration units, and null representations stay unchanged. LowCardinality(Time64) remains unsupported by the Rust core.

For nullable scalar DateTime64(9), `query_df` and `query_df_stream` preserve all nine fractional digits, including through SimpleAggregateFunction aliases. On Pandas 3, datetime columns in default output now use nanosecond units instead of microseconds. Existing object dtype inference for all-NULL blocks stays unchanged. If a named timezone would put a value outside Pandas' nanosecond range, that column keeps its previous object conversion and range behavior.

For these columns, `query_np` and `query_np_stream` object fields now contain `numpy.datetime64` cells instead of Python `datetime` cells, with `None` for SQL NULL. Named-zone values use UTC ticks without timezone metadata, matching the Python codec. This preserves nanoseconds without requiring Pandas. Use `query_df` if you need timezone-aware timestamps. Other nullable temporal NumPy cell types stay unchanged.

Scalar DateTime64 NumPy and Pandas queries reject precisions other than 0, 3, 6, and 9 with `ProgrammingError`, including nullable columns and SimpleAggregateFunction aliases.

If a Rust codec is selected and the compiled module is not installed, client creation raises a `NotSupportedError` naming this install command.

<h2 id="enabling-the-codec">
  Enabling the codec
</h2>

Select the codec with the `native_codec` client option:

```python theme={null}
import clickhouse_connect

client = clickhouse_connect.get_client(host="localhost", native_codec="rust")
```

Accepted values:

| Value | Behavior |
| - | - |
| `python` | Default. The existing Python and Cython codec. |
| `rust` | Prefer the Rust codec. Queries with unsupported options and inserts with unsupported types route to the Python codec. |
| `rust_strict` | Require the Rust codec. Unsupported options and types raise instead of routing. |

The default can also be seeded with the `native_codec` common setting or the `CLICKHOUSE_CONNECT_NATIVE_CODEC` environment variable. Precedence is the client keyword argument, then the common setting, then the environment variable.

The option is ignored for `interface="chdb"` clients, which always use the Python codec.

<h2 id="when-to-use-the-rust-codec">
  When to use the Rust codec
</h2>

The Rust codec is most useful for large DataFrame results with text, container, and complex types such as `String`, `LowCardinality`, `Map`, `Array`, `JSON`, `Decimal`, and `UUID`. It can also help large row and column block streams, concurrent query workloads, and bulk inserts.

Small results and network-bound queries may see little change. Flat numeric results already use efficient bulk paths in the Python codec, so they may also see less benefit.

Direct buffers avoid intermediate Arrow arrays for supported types. Final NumPy and Pandas result assembly can still allocate and copy, so this doesn't make every query result zero-copy.

Without PyArrow, Rust string conversion builds Python string objects from each decoded block. At default block sizes, peak memory for large String results is close to the Python codec. Very large blocks, such as a single block holding a million rows, keep the whole decoded block alive while its strings are built and can use more peak memory than the Python codec. For large Pandas `StringDtype` results, install the `arrow` extra and use Arrow string storage. Pandas 3 selects Arrow storage by default when PyArrow is installed. Pandas 2 needs `pd.options.mode.string_storage = "pyarrow"`.

Buffered `query()` calls on very wide or all-numeric results can currently be slower and use more peak memory with the Rust codec. Prefer `query_df`, `query_row_block_stream`, or `query_column_block_stream` for those workloads. For NumPy and Pandas, `query_np_stream` and `query_df_stream` can also reduce peak memory. Process and release each block instead of collecting the stream. Memory use still depends on the size of each block, so streaming a single large block doesn't provide the same benefit.

Benchmark your own workload before adopting the codec. Use `native_codec="rust_strict"` while measuring so an unsupported option or missing dependency raises instead of silently routing the query to Python.

<h2 id="fallback-rules">
  Fallback rules
</h2>

Fallback decisions are made before any bytes are read or sent, so there is never a mid-stream codec switch. For queries the choice happens before the response body is consumed. For inserts the Rust encoder is only selected when every column type is supported, otherwise the whole insert runs on the Python codec.

When `naive_datetime_insert="server"` is active, `rust` routes an insert containing any `DateTime` or `DateTime64` column to the Python codec so the declared column timezone or server timezone is applied. `rust_strict` rejects that combination. The default `naive_datetime_insert="local"` mode continues to use the Rust encoder.

Driver-internal metadata queries, including SQLAlchemy dialect reflection statements, always use the Python codec, silently, in every mode.

Malformed Native payloads detected by the Rust codec raise `DataError`.

<h2 id="versioning">
  Versioning
</h2>

`clickhouse-connect-core` versions independently of `clickhouse-connect`. The `rust` extra requires `clickhouse-connect-core>=0.2.1,<0.3`. The driver checks both the binding API version and the column-buffer capability when a Rust codec is selected. The published core 0.2.0 wheel lacks the column-buffer capability. If the installed wheel is incompatible, client creation raises a `NotSupportedError` for both `rust` and `rust_strict`, with this upgrade command:

```bash theme={null}
pip install --upgrade clickhouse-connect-core
```

Fixes and performance improvements in the compiled codec ship as `clickhouse-connect-core` releases and can be picked up with a compatible core wheel upgrade. Changes to the driver's Rust integration require a `clickhouse-connect` upgrade.

<h2 id="known-behavior-differences">
  Known behavior differences
</h2>

The Rust codec targets cell for cell parity with the Python codec. The following differences are known.

* `query_np` and `query_df` results for `Variant` columns contain plain Python objects rather than numpy scalar values. The values are equal, the cell types differ.
* `Dynamic` values that contain `Time64` materialize as `datetime.timedelta` in Rust `query_np` and `query_df` results. At scales 0, 3, 6, and 9 the values equal the Python codec's `numpy.timedelta64` cells, but the cell types differ. At other scales the Rust codec returns `datetime.timedelta` while the Python codec raises `ProgrammingError` because NumPy has no matching unit. Dynamic member metadata is not exposed to the driver after Rust decoding, so use `native_codec="python"` when NumPy cell types or unsupported-scale validation are required.
* For `query_df`, the Python codec may stringify compound values stored in JSON shared data. The Rust codec returns decoded objects, matching both codecs' `query_np` results.
* A `LowCardinality` alternative inside a container that materializes per cell, such as `Array(Variant(...))`, produces value-equal cells that do not share the per-dictionary-slot object identity the Python codec exhibits.
* `Nullable(Tuple(...))` columns with one or more elements decode correctly on the Rust codec. The Python codec misreads this layout and the Rust result is the reference behavior. Both codecs support `Nullable(Tuple())`.
* `rust_strict` rejects query options the Rust path does not implement, such as custom per-query `query_formats`, rather than silently changing behavior.
* Rust insert conversion and validation errors can raise `DataError` where the Python codec raises `ValueError`, and the message text can differ. Examples include invalid `Time` and `Time64` values, IPv6 addresses, QBit dimensions, `FixedString` lengths, `Float64` strings, and attempts to insert elements into a `Tuple()` column.
* The Rust encoder rejects `b""` for `FixedString(N)` and numeric strings such as `"2"` for integer columns. The Python codec zero-fills an empty `FixedString` value and coerces numeric strings.
* The Rust codec decodes some `Dynamic` shared-variant values to their Python types when the Python codec leaves the value as raw bytes. For example, a stored `Decimal` can return as `decimal.Decimal` from Rust and as a binary value from Python.
* For `Date` values in `Dynamic` shared storage, `query_np` and `query_df` return `datetime.date` cells with the Rust codec and `numpy.datetime64` cells with the Python codec. Both codecs return `datetime.date` from `query`.
