Dialects#
SQLSpec registers custom sqlglot dialects that extend built-in SQL grammars with extension-specific operators. These dialects enable the builder to parse and generate SQL that uses vendor-specific syntax (e.g., pgvector distance operators, ParadeDB search operators).
The dialects are registered lazily through the sqlglot.dialects entry-point
group, so sqlglot.parse_one(..., dialect="pgvector") resolves them in any
environment where SQLSpec is installed; importing sqlspec alone does not load
them. Import the classes directly when you need them as objects:
from sqlspec.dialects import PGTextSearch, PGVector, ParadeDB, Spanner, Spangres
Performance builds compile the custom dialect helper modules alongside
sqlglot[c]: generator transforms, operator registries, and compatibility
helpers. The small SQLGlot subclass/registration modules stay interpreted
because native-compiling those tokenizer/dialect classes is not runtime-safe.
PostgreSQL Extensions#
PGVector#
- class sqlspec.dialects.postgres.PGVector[source]
Bases:
PostgresPostgreSQL dialect with pgvector extension support.
- Tokenizer
alias of
PGVectorTokenizer
- BYTE_STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = False
Whether byte string literals support escape sequences. Set by the metaclass based on the tokenizer's BYTE_STRING_ESCAPES.
- STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = False
Whether string literals support escape sequences (e.g. n). Set by the metaclass based on the tokenizer's STRING_ESCAPES.
- SUPPORTS_COLUMN_JOIN_MARKS = False
Whether the old-style outer join (+) syntax is supported.
- tokenizer_class
alias of
PGVectorTokenizer
Adds support for pgvector distance operators:
Operator |
Description |
|---|---|
|
L2 (Euclidean) distance |
|
Negative inner product |
|
Cosine distance |
|
L1 (Manhattan) distance |
|
Hamming distance (binary vectors) |
|
Jaccard distance (binary vectors) |
PGTextSearch#
- class sqlspec.dialects.postgres.PGTextSearch[source]
Bases:
PostgresPostgreSQL dialect with pg_textsearch and pgvector extension support.
- Tokenizer
alias of
PGTextSearchTokenizer
- BYTE_STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = False
Whether byte string literals support escape sequences. Set by the metaclass based on the tokenizer's BYTE_STRING_ESCAPES.
- STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = False
Whether string literals support escape sequences (e.g. n). Set by the metaclass based on the tokenizer's STRING_ESCAPES.
- SUPPORTS_COLUMN_JOIN_MARKS = False
Whether the old-style outer join (+) syntax is supported.
- tokenizer_class
alias of
PGTextSearchTokenizer
Adds support for PostgreSQL deployments with the pg_textsearch BM25 extension installed.
AlloyDB currently provides this extension in preview on PostgreSQL 17 and 18; see
AlloyDB BM25 requirements.
Server packaging and version support are independent of SQLSpec dialect availability.
Operator |
Description |
|---|---|
|
BM25 score ranking operator (returns negative score for ASC index scans) |
Indexes are created with USING bm25 (column) WITH (text_config='english') and queries order by column <@> 'query' ASC.
Enable the extension in the database before using its operators or index method.
The dialect also supports pgvector distance operators for hybrid queries.
Asyncpg, psycopg, psqlpy, and PostgreSQL-backed ADBC configurations probe enabled
extensions on first connection. enable_pg_textsearch defaults to True;
setting it to False disables detection, not the installed server extension.
pg_textsearch_available and is_postgres_extension_active() report the cached,
enabled-and-detected state, and remain false before a successful probe.
ADK enable_bm25 requires successful pg_textsearch detection.
The dialect label remains selected in priority order: paradedb,
pg_textsearch, then pgvector for an otherwise default PostgreSQL
configuration. The active extension set records all enabled discoveries independently.
An explicitly selected non-default dialect is preserved; select a dialect that
supports the operators your queries use. Extension detection does not override
that choice, and a dialect label alone does not mark an extension as available.
ParadeDB#
- class sqlspec.dialects.postgres.ParadeDB[source]
Bases:
PostgresParadeDB dialect with pg_search and pgvector extension support.
- Tokenizer
alias of
ParadeDBTokenizer
- BYTE_STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = False
Whether byte string literals support escape sequences. Set by the metaclass based on the tokenizer's BYTE_STRING_ESCAPES.
- STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = False
Whether string literals support escape sequences (e.g. n). Set by the metaclass based on the tokenizer's STRING_ESCAPES.
- SUPPORTS_COLUMN_JOIN_MARKS = False
Whether the old-style outer join (+) syntax is supported.
- tokenizer_class
alias of
ParadeDBTokenizer
Extends PostgreSQL with ParadeDB (pg_search) operators:
Operator |
Description |
|---|---|
|
BM25 full-text search / complex query expressions |
|
Conjunction match (all tokenized terms must match) |
|
Disjunction match (any tokenized term matches) |
|
Exact term match (no tokenization of right-hand side) |
|
Exact phrase match (token order and position enforced) |
|
Proximity match in any order ( |
|
Ordered proximity match (left term must appear first) |
Scoring and snippets in ParadeDB are standard SQL functions (pdb.score(),
pdb.snippet()), so they require no custom operator syntax.
Spanner#
- class sqlspec.dialects.spanner.Spanner[source]
Bases:
BigQueryGoogle Cloud Spanner SQL dialect.
- Tokenizer
alias of
SpannerTokenizer
- parse(sql, **opts)[source]
Parse Spanner SQL statements and attach hints.
- parse_into(expression_type, sql, **opts)[source]
Parse into specific expression type with attached hints.
- BYTE_STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = True
Whether byte string literals support escape sequences. Set by the metaclass based on the tokenizer's BYTE_STRING_ESCAPES.
- STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = True
Whether string literals support escape sequences (e.g. n). Set by the metaclass based on the tokenizer's STRING_ESCAPES.
- SUPPORTS_COLUMN_JOIN_MARKS = False
Whether the old-style outer join (+) syntax is supported.
- UNESCAPED_SEQUENCES: dict[str, str] = {'\\\\': '\\', '\\a': '\x07', '\\b': '\x08', '\\f': '\x0c', '\\n': '\n', '\\r': '\r', '\\t': '\t', '\\v': '\x0b'}
Mapping of an escaped sequence (n) to its unescaped version (` `).
- tokenizer_class
alias of
SpannerTokenizer
- class sqlspec.dialects.spanner.Spangres[source]
Bases:
PostgresSpanner PostgreSQL-compatible dialect.
- parse(sql, **opts)[source]
Parse Spangres SQL statements and attach hints.
- parse_into(expression_type, sql, **opts)[source]
Parse into specific expression type with attached hints.
- BYTE_STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = False
Whether byte string literals support escape sequences. Set by the metaclass based on the tokenizer's BYTE_STRING_ESCAPES.
- STRINGS_SUPPORT_ESCAPED_SEQUENCES: bool = False
Whether string literals support escape sequences (e.g. n). Set by the metaclass based on the tokenizer's STRING_ESCAPES.
- SUPPORTS_COLUMN_JOIN_MARKS = False
Whether the old-style outer join (+) syntax is supported.
Expression Types#
- sqlspec.builder.VectorDistance(*, this, expression, metric='euclidean')[source]
Build a SQLSpec vector-distance expression.
- Return type:
Operator