FORMAT Native traffic, which covers query, query_np, query_df, their block and row streaming variants, and inserts including insert_df. The Arrow methods use FORMAT Arrow and are unaffected, as are raw queries, raw inserts, and non-Native formats.
The codec is experimental and opt in. The Python codec remains the default.
Installation
The compiled codec ships as a separate wheel namedclickhouse-connect-core, which provides the _ch_core extension module. Install the codec with the integrations your application uses:
query_np, query_df, their streaming variants, and query(..., use_numpy=True). NumPy queries don’t require Pandas. Add the async extra for the async client.
PyArrow is still required for Arrow methods and Arrow-backed Pandas storage. Install the arrow extra for those features:
mode.string_storage option. When PyArrow is installed, strings use the existing Arrow conversion, with full UTF-8 validation, for either storage backend. Without PyArrow, the driver uses the Rust object conversion and Pandas’ default storage uses Python strings. Explicitly requesting Arrow string storage without PyArrow raises a dependency error. Invalid UTF-8 keeps the existing hex rendering with either storage backend.
The codec, its packaging, and its dependency set are experimental. Direct buffer conversion now handles non-nullable 8-64-bit integers, Float32/64, Boolean, BFloat16, Interval, Date, Date32, DateTime, and DateTime64 columns. Scalar Time and Time64 columns, including nullable values, also use direct buffers, as do nullable BFloat16 and extended Pandas nullable integer, float, and Interval output. BFloat16 retains float32 output and its existing Pandas Float32 missing-value policy. Interval values remain signed counts in their declared unit. Temporal conversion retains declared units and existing timezone rules. Nullable SimpleAggregateFunction aliases of Time64 now return correct durations and NaT for NULL values. Nullable scalar DateTime64(9) columns also use buffers to preserve nanoseconds in NumPy and Pandas output. Other nullable dates and timestamps retain their object conversion paths.
Arrays of Time/Time64, including nested arrays and nullable leaves, and LowCardinality(Time), including nullable values, also use direct buffers. Their nesting, duration units, and null representations stay unchanged. LowCardinality(Time64) remains unsupported by the Rust core.
For nullable scalar DateTime64(9), query_df and query_df_stream preserve all nine fractional digits, including through SimpleAggregateFunction aliases. On Pandas 3, datetime columns in default output now use nanosecond units instead of microseconds. Existing object dtype inference for all-NULL blocks stays unchanged. If a named timezone would put a value outside Pandas’ nanosecond range, that column keeps its previous object conversion and range behavior.
For these columns, query_np and query_np_stream object fields now contain numpy.datetime64 cells instead of Python datetime cells, with None for SQL NULL. Named-zone values use UTC ticks without timezone metadata, matching the Python codec. This preserves nanoseconds without requiring Pandas. Use query_df if you need timezone-aware timestamps. Other nullable temporal NumPy cell types stay unchanged.
Scalar DateTime64 NumPy and Pandas queries reject precisions other than 0, 3, 6, and 9 with ProgrammingError, including nullable columns and SimpleAggregateFunction aliases.
If a Rust codec is selected and the compiled module is not installed, client creation raises a NotSupportedError naming this install command.
Enabling the codec
Select the codec with thenative_codec client option:
The default can also be seeded with the
native_codec common setting or the CLICKHOUSE_CONNECT_NATIVE_CODEC environment variable. Precedence is the client keyword argument, then the common setting, then the environment variable.
The option is ignored for interface="chdb" clients, which always use the Python codec.
When to use the Rust codec
The Rust codec is most useful for large DataFrame results with text, container, and complex types such asString, LowCardinality, Map, Array, JSON, Decimal, and UUID. It can also help large row and column block streams, concurrent query workloads, and bulk inserts.
Small results and network-bound queries may see little change. Flat numeric results already use efficient bulk paths in the Python codec, so they may also see less benefit.
Direct buffers avoid intermediate Arrow arrays for supported types. Final NumPy and Pandas result assembly can still allocate and copy, so this doesn’t make every query result zero-copy.
Without PyArrow, Rust string conversion builds Python string objects from each decoded block. At default block sizes, peak memory for large String results is close to the Python codec. Very large blocks, such as a single block holding a million rows, keep the whole decoded block alive while its strings are built and can use more peak memory than the Python codec. For large Pandas StringDtype results, install the arrow extra and use Arrow string storage. Pandas 3 selects Arrow storage by default when PyArrow is installed. Pandas 2 needs pd.options.mode.string_storage = "pyarrow".
Buffered query() calls on very wide or all-numeric results can currently be slower and use more peak memory with the Rust codec. Prefer query_df, query_row_block_stream, or query_column_block_stream for those workloads. For NumPy and Pandas, query_np_stream and query_df_stream can also reduce peak memory. Process and release each block instead of collecting the stream. Memory use still depends on the size of each block, so streaming a single large block doesn’t provide the same benefit.
Benchmark your own workload before adopting the codec. Use native_codec="rust_strict" while measuring so an unsupported option or missing dependency raises instead of silently routing the query to Python.
Fallback rules
Fallback decisions are made before any bytes are read or sent, so there is never a mid-stream codec switch. For queries the choice happens before the response body is consumed. For inserts the Rust encoder is only selected when every column type is supported, otherwise the whole insert runs on the Python codec. Whennaive_datetime_insert="server" is active, rust routes an insert containing any DateTime or DateTime64 column to the Python codec so the declared column timezone or server timezone is applied. rust_strict rejects that combination. The default naive_datetime_insert="local" mode continues to use the Rust encoder.
Driver-internal metadata queries, including SQLAlchemy dialect reflection statements, always use the Python codec, silently, in every mode.
Malformed Native payloads detected by the Rust codec raise DataError.
Versioning
clickhouse-connect-core versions independently of clickhouse-connect. The rust extra requires clickhouse-connect-core>=0.2.1,<0.3. The driver checks both the binding API version and the column-buffer capability when a Rust codec is selected. The published core 0.2.0 wheel lacks the column-buffer capability. If the installed wheel is incompatible, client creation raises a NotSupportedError for both rust and rust_strict, with this upgrade command:
clickhouse-connect-core releases and can be picked up with a compatible core wheel upgrade. Changes to the driver’s Rust integration require a clickhouse-connect upgrade.
Known behavior differences
The Rust codec targets cell for cell parity with the Python codec. The following differences are known.query_npandquery_dfresults forVariantcolumns contain plain Python objects rather than numpy scalar values. The values are equal, the cell types differ.Dynamicvalues that containTime64materialize asdatetime.timedeltain Rustquery_npandquery_dfresults. At scales 0, 3, 6, and 9 the values equal the Python codec’snumpy.timedelta64cells, but the cell types differ. At other scales the Rust codec returnsdatetime.timedeltawhile the Python codec raisesProgrammingErrorbecause NumPy has no matching unit. Dynamic member metadata is not exposed to the driver after Rust decoding, so usenative_codec="python"when NumPy cell types or unsupported-scale validation are required.- For
query_df, the Python codec may stringify compound values stored in JSON shared data. The Rust codec returns decoded objects, matching both codecs’query_npresults. - A
LowCardinalityalternative inside a container that materializes per cell, such asArray(Variant(...)), produces value-equal cells that do not share the per-dictionary-slot object identity the Python codec exhibits. Nullable(Tuple(...))columns with one or more elements decode correctly on the Rust codec. The Python codec misreads this layout and the Rust result is the reference behavior. Both codecs supportNullable(Tuple()).rust_strictrejects query options the Rust path does not implement, such as custom per-queryquery_formats, rather than silently changing behavior.- Rust insert conversion and validation errors can raise
DataErrorwhere the Python codec raisesValueError, and the message text can differ. Examples include invalidTimeandTime64values, IPv6 addresses, QBit dimensions,FixedStringlengths,Float64strings, and attempts to insert elements into aTuple()column. - The Rust encoder rejects
b""forFixedString(N)and numeric strings such as"2"for integer columns. The Python codec zero-fills an emptyFixedStringvalue and coerces numeric strings. - The Rust codec decodes some
Dynamicshared-variant values to their Python types when the Python codec leaves the value as raw bytes. For example, a storedDecimalcan return asdecimal.Decimalfrom Rust and as a binary value from Python. - For
Datevalues inDynamicshared storage,query_npandquery_dfreturndatetime.datecells with the Rust codec andnumpy.datetime64cells with the Python codec. Both codecs returndatetime.datefromquery.