Expose a catalog
A worker does not have to be a bag of functions. It can present itself as a database: a named
catalog you ATTACH, holding schemas that hold tables and views, queried with ordinary qualified
names. The user never sees a function call.
The whole worker
Section titled “The whole worker”catalog.rs
// Copyright 2025, 2026 Query Farm LLC - https://query.farm
//! The catalog example for the vgi-rust documentation.
//!
//! A worker does not have to be a bag of functions. It can present itself as a
//! database: a named catalog you ATTACH, holding schemas that hold tables and
//! views, queried with ordinary qualified names.
//!
//! The table here is *function-backed*: `CatTable` names a table function plus
//! the arguments to call it with, so `SELECT * FROM cat.data.cities` runs the
//! function with those arguments baked in. The user never passes them, and
//! never sees the function.
//!
//! ```text
//! cargo build --release --bin catalog
//! # then, in a Haybarn shell:
//! ATTACH 'cat' (TYPE vgi, LOCATION './target/release/catalog');
//! SELECT * FROM cat.data.cities;
//! SELECT * FROM cat.data.big_cities; -- a view over the table
//! ```
use std::sync::Arc;
use arrow_array::{ArrayRef, Int64Array, RecordBatch, StringArray};
use arrow_schema::{DataType, Field, Schema, SchemaRef};
use vgi::arguments::Arguments;
use vgi::catalog::{CatSchema, CatTable, CatView, CatalogModel};
use vgi::function::{ArgSpec, BindParams, BindResponse, FunctionMetadata, ProcessParams};
use vgi::table_function::{TableFunction, TableProducer};
use vgi::vgi_rpc::OutputCollector;
use vgi::{Result, RpcError};
/// The table's shape. Declaring it on the CatTable lets DuckDB describe the
/// table without calling the worker at all.
fn cities_schema() -> SchemaRef {
Arc::new(Schema::new(vec![
Field::new("name", DataType::Utf8, true),
Field::new("population", DataType::Int64, true),
]))
}
/// Stands in for whatever the worker actually fronts — a remote API, a file
/// format, a device.
const CITIES: &[(&str, i64)] = &[
("Charlottesville", 51_000),
("Richmond", 230_000),
("Virginia Beach", 457_000),
];
const BATCH_SIZE: usize = 1024;
/// The scan cursor.
///
/// It holds an *index*, not the rows. That matters: a producer is carried
/// between `next_batch` calls and may be serialized for an HTTP continuation,
/// so materializing the whole result here would be paid for repeatedly. Three
/// rows would not notice; three million would.
struct CitiesProducer {
schema: SchemaRef,
min_population: i64,
next: usize,
}
impl TableProducer for CitiesProducer {
fn next_batch(&mut self, _out: &mut OutputCollector) -> Result<Option<RecordBatch>> {
let mut names: Vec<&str> = Vec::new();
let mut pops: Vec<i64> = Vec::new();
while self.next < CITIES.len() && names.len() < BATCH_SIZE {
let (name, pop) = CITIES[self.next];
self.next += 1;
if pop >= self.min_population {
names.push(name);
pops.push(pop);
}
}
if names.is_empty() {
return Ok(None);
}
let cols: Vec<ArrayRef> = vec![
Arc::new(StringArray::from(names)),
Arc::new(Int64Array::from(pops)),
];
RecordBatch::try_new(self.schema.clone(), cols)
.map(Some)
.map_err(|e| RpcError::runtime_error(e.to_string()))
}
}
/// The scan behind the table. It is an ordinary table function — nothing about
/// it knows it is backing a catalog table.
struct CitiesScan;
impl TableFunction for CitiesScan {
fn name(&self) -> &str {
"cities_scan"
}
fn metadata(&self) -> FunctionMetadata {
FunctionMetadata {
description: "Scans the cities table, optionally filtered by population".to_string(),
..Default::default()
}
}
fn argument_specs(&self) -> Vec<ArgSpec> {
vec![ArgSpec::const_arg(
"min_population",
0,
"int64",
"Only cities at least this large",
)
.with_ge(0.0)]
}
fn on_bind(&self, _params: &BindParams) -> Result<BindResponse> {
Ok(BindResponse {
output_schema: cities_schema(),
opaque_data: Vec::new(),
})
}
fn producer(&self, params: &ProcessParams) -> Result<Box<dyn TableProducer>> {
Ok(Box::new(CitiesProducer {
schema: params.output_schema.clone(),
min_population: params.arguments.const_i64(0).unwrap_or(0),
next: 0,
}))
}
}
fn main() {
let mut worker = vgi::Worker::new();
worker.register_table(CitiesScan);
let cities = CatTable {
name: "cities".to_string(),
columns: cities_schema(),
// Function-backed: the scan function plus the arguments to call it with.
scan_function: "cities_scan".to_string(),
scan_arguments: Arguments::serialize_scan_args(&[Arc::new(Int64Array::from(vec![0i64]))])
.unwrap_or_default(),
comment: Some("Every city the worker knows about".to_string()),
// Column 0 (`name`) is NOT NULL — surfaced in duckdb_columns().
not_null: vec![0],
..Default::default()
};
let big_cities = CatView {
name: "big_cities".to_string(),
// Pure SQL that DuckDB evaluates — no worker round trip for the view
// itself, only for the table it reads.
//
// `cat` is hardcoded, which would normally be fragile. It is not: the
// name in ATTACH must match this catalog's name, so `cat` is the only
// name this view is ever read under.
definition: "SELECT * FROM cat.data.cities WHERE population >= 100000".to_string(),
comment: Some("Cities with a population of at least 100,000".to_string()),
..Default::default()
};
worker.set_catalog(CatalogModel {
name: "cat".to_string(),
comment: Some("Documentation example: a worker presented as a database".to_string()),
schemas: vec![CatSchema {
name: "data".to_string(),
tables: vec![cities],
views: vec![big_cities],
..Default::default()
}],
..Default::default()
});
worker.run();
}
ATTACH 'cat' (TYPE vgi, LOCATION './target/release/catalog');
SELECT * FROM cat.data.cities;
Output
| name | population |
|---|---|
| Charlottesville | 51000 |
| Richmond | 230000 |
| Virginia Beach | 457000 |
Function-backed tables
Section titled “Function-backed tables”The table is not stored — it is a table function plus the arguments to call it with:
let cities = CatTable {
name: "cities".to_string(),
columns: cities_schema(),
scan_function: "cities_scan".to_string(),
scan_arguments: Arguments::serialize_scan_args(&[Arc::new(Int64Array::from(vec![0i64]))])
.unwrap_or_default(),
..Default::default()
};
SELECT * FROM cat.data.cities runs cities_scan(0). The user passes nothing and never learns the
function exists. Declaring columns explicitly is what lets DuckDB describe the table without
calling the worker at all.
scan_arguments is an IPC-serialized argument batch rather than a Rust list, because it travels on
the wire in the catalog description; Arguments::serialize_scan_args is the encoder.
VGI tables live in DuckDB’s catalog model, so you query them with ordinary qualified names, and
duckdb_tables() and duckdb_columns() see them. The difference is where the rows come from: a
native table is stored and managed by DuckDB, while a VGI table delegates its scan to your worker.
A view is pure SQL. Once its definition has been advertised at attach time, DuckDB evaluates it itself — there is no worker round trip for the view, only for whatever tables it reads.
let big_cities = CatView {
name: "big_cities".to_string(),
definition: "SELECT * FROM cat.data.cities WHERE population >= 100000".to_string(),
..Default::default()
};
Output
| name | population |
|---|---|
| Richmond | 230000 |
| Virginia Beach | 457000 |
cat is hardcoded there, which normally would be fragile — but the name in ATTACH is not a free
alias. It must match CatalogModel.name, and attaching under any other name fails at once with
“No worker handles catalog ‘mydb’”. So the catalog name your worker declares is the name the view
will always be read under.
Metadata surfaces
Section titled “Metadata surfaces”Constraints and comments are not decoration — they reach DuckDB’s own catalog views:
SELECT column_name, is_nullable FROM duckdb_columns() WHERE table_name = 'cities';
Output
| column_name | is_nullable |
|---|---|
| name | false |
| population | true |
not_null: vec![0] became is_nullable = false on column 0. Note that the constraint fields take
column indices, not names — a rename that does not update them moves the constraint silently.
What else a table can declare
Section titled “What else a table can declare”CatTable carries considerably more than the example uses. The ones worth knowing early:
| Field | What it does |
|---|---|
not_null, unique, primary_key, check, foreign_keys | Constraints, surfaced to the planner and to duckdb_constraints(). |
cardinality | A row-count estimate DuckDB can plan against without calling the worker. |
statistics | Per-column min/max, so the optimizer can eliminate impossible filters at plan time. |
required_filters | Refuse an unbounded scan. An AND of OR-groups of column paths; the extension throws a BinderException naming any unsatisfied group. |
supports_time_travel, time_travel, branches | Answer AT (VERSION …) / AT (TIMESTAMP …) queries. |
scan_function_impl | Attach the function value directly rather than resolving it by name. |
The example’s producer holds an index, not the rows. That matters: a producer is carried between
next_batch calls and may be serialized for an HTTP continuation, so materializing the whole result
in it is paid for repeatedly. Three rows would not notice; three million would.
Next steps
Section titled “Next steps”- Making the scan smarter → Function patterns.
- Serving it remotely → Transports & runtimes.
- Exact types → Catalogs.