<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
<title>CROWDB Blog</title><subtitle>Notes from inside a distributed storage system</subtitle>
<link href="https://buzzcrow.github.io/feed.xml" rel="self"/><link href="https://buzzcrow.github.io/"/>
<id>https://buzzcrow.github.io/</id><updated>2026-09-30T18:04:48+00:00</updated><author><name>Gian Crow</name></author>
<entry><title>Load TPC-H and TPC-DS tables into CROWDB Iceberg</title><link href="https://buzzcrow.github.io/blog/load-tpc-h-and-tpc-ds-into-crowdb-iceberg/"/><id>https://buzzcrow.github.io/blog/load-tpc-h-and-tpc-ds-into-crowdb-iceberg/</id><published>2026-09-30T15:20:00+00:00</published><updated>2026-09-30T18:01:00+00:00</updated><summary>Load 8 TPC-H or 24 TPC-DS tables with one command, then select rows from a fresh Python client.</summary><content type="html">&lt;p&gt;I wanted a repeatable way to put recognizable data into the CROWDB Iceberg container before trying SQL engines. Hand-writing a sample table proves little about multi-table imports. &lt;a href=&quot;https://github.com/buzzcrow/crowdb-tpc-loader&quot;&gt;crowdb-tpc-loader&lt;/a&gt; now generates TPC-H and TPC-DS Parquet, checks every table, uploads the files, and registers them through the Iceberg REST Catalog. This guide uses scale factor 0.01 to exercise that path. It does not run the benchmark queries or claim TPC certification.&lt;/p&gt;

&lt;h2 id=&quot;start-a-local-container&quot;&gt;Start a local container&lt;/h2&gt;

&lt;p&gt;These commands use a Linux amd64 development checkout of &lt;a href=&quot;https://github.com/buzzcrow/crowdb&quot;&gt;CROWDB&lt;/a&gt; at version 0.2.0. The image below is &lt;strong&gt;built locally from source&lt;/strong&gt;; the test does not establish that a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;0.2.0&lt;/code&gt; Docker Hub tag has been published. The single-node profile has one copy and no node-failure protection, so use disposable data.&lt;/p&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;cd&lt;/span&gt; /path/to/crowdb
pixi run build-single-node-container
docker run &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--name&lt;/span&gt; crowdb-tpc-demo &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; 127.0.0.1:80:80 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;-v&lt;/span&gt; crowdb-tpc-demo-data:/opt/crowdb/data &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  crowdb-iceberg-single-node:dev
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Wait until &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;docker inspect crowdb-tpc-demo --format &apos;&apos;&lt;/code&gt; says &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;healthy&lt;/code&gt;. Port 80 carries the REST Catalog and native Iceberg file endpoint. Get the local client credentials without copying the token into a command line:&lt;/p&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;umask &lt;/span&gt;077
docker &lt;span class=&quot;nb&quot;&gt;exec &lt;/span&gt;crowdb-tpc-demo crowdb-monitor credentials show &lt;span class=&quot;nt&quot;&gt;--format&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;env&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; /tmp/crowdb-tpc-demo.env
&lt;span class=&quot;nb&quot;&gt;set&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;-a&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;.&lt;/span&gt; /tmp/crowdb-tpc-demo.env
&lt;span class=&quot;nb&quot;&gt;set&lt;/span&gt; +a
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Keep that file private and remove it when finished. Use a different host port or adjust &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ICEBERG_URI&lt;/code&gt; if port 80 is occupied.&lt;/p&gt;

&lt;h2 id=&quot;load-both-datasets&quot;&gt;Load both datasets&lt;/h2&gt;

&lt;p&gt;From a checkout of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;crowdb-tpc-loader&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;cd&lt;/span&gt; /path/to/crowdb-tpc-loader
python3 &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; venv .venv
&lt;span class=&quot;nb&quot;&gt;.&lt;/span&gt; .venv/bin/activate
python &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; pip &lt;span class=&quot;nb&quot;&gt;install&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--only-binary&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;:all: &lt;span class=&quot;nt&quot;&gt;-e&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;.&lt;/span&gt;

crowdb-tpc-loader load &lt;span class=&quot;nt&quot;&gt;--benchmark&lt;/span&gt; tpch &lt;span class=&quot;nt&quot;&gt;--sf&lt;/span&gt; 0.01 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--namespace&lt;/span&gt; tpch_demo &lt;span class=&quot;nt&quot;&gt;--report-file&lt;/span&gt; ./tpch-demo.json
crowdb-tpc-loader load &lt;span class=&quot;nt&quot;&gt;--benchmark&lt;/span&gt; tpcds &lt;span class=&quot;nt&quot;&gt;--sf&lt;/span&gt; 0.01 &lt;span class=&quot;se&quot;&gt;\&lt;/span&gt;
  &lt;span class=&quot;nt&quot;&gt;--namespace&lt;/span&gt; tpcds_demo &lt;span class=&quot;nt&quot;&gt;--report-file&lt;/span&gt; ./tpcds-demo.json
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The commands need a fresh namespace. TPC-H uses the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tpchgen-cli&lt;/code&gt; 3.0.0 binary; TPC-DS uses DuckDB 1.5’s official &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tpcds&lt;/code&gt; extension. The first run may download these generator components. The loader validates the full generated dataset before creating a benchmark table, then uploads each Parquet file and registers it with Iceberg. CROWDB checks the upload’s signed payload or supplied checksum before accepting it; the loader does not download the object for another checksum pass. The loader uploads up to 24 files concurrently by default, then commits one snapshot per table in order. It deletes local staging on success; the JSON reports retain remote file locations, counts, and snapshot IDs. An existing table stops a default load, so choose a new namespace when repeating the experiment.&lt;/p&gt;

&lt;p&gt;In my local SF 0.01 run, TPC-H produced &lt;strong&gt;8 tables, 86,805 rows, and 3.23 MB across 8 Parquet files&lt;/strong&gt;. TPC-DS produced &lt;strong&gt;24 tables, 277,976 rows, and 3.86 MB across 24 Parquet files&lt;/strong&gt;. There was one file per table in this small run, but the loader accepts multiple shards. TPC-H &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;region&lt;/code&gt; was 1,227 bytes while &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lineitem&lt;/code&gt; was 1,924,571 bytes; TPC-DS &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;call_center&lt;/code&gt; was 4,917 bytes while &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;store_sales&lt;/code&gt; was 985,882 bytes. These are compressed Parquet sizes, not in-memory table sizes.&lt;/p&gt;

&lt;h2 id=&quot;select-rows-in-python&quot;&gt;Select rows in Python&lt;/h2&gt;

&lt;p&gt;Run this in a fresh process after the loader has removed its local files. The file adapter handles exact-object requests on CROWDB’s native Iceberg endpoint:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;from&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pyiceberg.catalog&lt;/span&gt; &lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;load_catalog&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;catalog&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;load_catalog&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
    &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;crowdb&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;type&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;rest&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;uri&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;environ&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;ICEBERG_URI&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;token&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;environ&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;ICEBERG_TOKEN&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;
    &lt;span class=&quot;o&quot;&gt;**&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;py-io-impl&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;crowdb_tpc_loader.crowdb_fileio.CrowdbFileIO&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;},&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;region&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;catalog&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;load_table&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;tpch_demo.region&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;nf&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;region&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;scan&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;row_filter&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;r_regionkey == 1&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;selected_fields&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;r_regionkey&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;r_name&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;to_arrow&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;().&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;to_pylist&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;())&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;item&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;catalog&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;load_table&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;tpcds_demo.item&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;nf&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;item&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;scan&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;selected_fields&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;i_item_sk&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;i_item_id&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;limit&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;5&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;to_arrow&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;().&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;to_pylist&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The first selection returned &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;{&apos;r_regionkey&apos;: 1, &apos;r_name&apos;: &apos;AMERICA&apos;}&lt;/code&gt; in my run. The second returned five &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;item&lt;/code&gt; rows. For a read-only check of every imported table, including remote Parquet footers and an Iceberg sample scan, run &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python scripts/verify_crowdb.py ./tpch-demo.json --require-complete --iceberg-scan&lt;/code&gt; and the same command with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tpcds-demo.json&lt;/code&gt;.&lt;/p&gt;

&lt;h2 id=&quot;query-from-duckdb&quot;&gt;Query from DuckDB&lt;/h2&gt;

&lt;p&gt;I also tested the locally built DuckDB 1.5.6 CLI with its &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;iceberg&lt;/code&gt; extension against the CROWDB REST Catalog. With the same credential environment loaded, this selected &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;AMERICA&lt;/code&gt; from the imported &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;region&lt;/code&gt; table:&lt;/p&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;/path/to/duckdb/build/release/duckdb &lt;span class=&quot;o&quot;&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;no&quot;&gt;SQL&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;
INSTALL iceberg;
LOAD iceberg;
INSTALL httpfs;
LOAD httpfs;
CREATE SECRET crowdb_catalog (TYPE ICEBERG, TOKEN &apos;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$ICEBERG_TOKEN&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;);
ATTACH &apos;&apos; AS crowdb (TYPE ICEBERG, SECRET crowdb_catalog, ENDPOINT &apos;&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;$ICEBERG_URI&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&apos;);
SELECT r_regionkey, r_name FROM crowdb.tpch_demo.region WHERE r_regionkey = 1;
&lt;/span&gt;&lt;span class=&quot;no&quot;&gt;SQL
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The empty &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ATTACH&lt;/code&gt; warehouse selector matters: this CROWDB profile has no named warehouse. DuckDB must use the object credentials supplied by the catalog for this Iceberg table. A separate fixed S3 credential from the monitor could attach the catalog but got HTTP 403 when reading the table’s manifest. I also queried the larger TPC-H SF 10 &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lineitem&lt;/code&gt; file and the newly loaded TPC-DS SF 10 &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;item&lt;/code&gt; table; the latter returned &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;(1, &apos;AAAAAAAABAAAAAAA&apos;)&lt;/code&gt; for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;i_item_sk, i_item_id&lt;/code&gt;. These are spot queries through the REST Catalog and file endpoint, not TPC query coverage or performance results.&lt;/p&gt;

&lt;p&gt;CROWDB’s Iceberg file paths look like S3 paths, but they are native Iceberg objects with table-scoped authority. &lt;a href=&quot;https://py.iceberg.apache.org/reference/pyiceberg/io/pyarrow/&quot;&gt;PyIceberg’s default PyArrow FileIO&lt;/a&gt; calls &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;get_file_info&lt;/code&gt; for an exact file; in this run that led to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ListObjectsV2&lt;/code&gt; and failed. The loader’s small FileIO adapter uses an exact-object request for existence and size. &lt;a href=&quot;https://iceberg.apache.org/docs/latest/fileio/&quot;&gt;Iceberg’s FileIO guide&lt;/a&gt; describes read, write, and seek as essential file operations. General prefix listing and the meaning of the S3 bucket field remain separate design work; this import does not require either.&lt;/p&gt;

&lt;h2 id=&quot;what-the-small-run-measured&quot;&gt;What the small run measured&lt;/h2&gt;

&lt;p&gt;I ran one single-node container on a Linux x86_64 host with an Intel Core i9-7960X, 32 logical CPUs, and 62 GiB RAM. The generator used two threads and a 1 GB DuckDB memory limit; Python 3.12.3, PyArrow 23.0.1, and PyIceberg 0.10.0 were installed. There was no performance baseline or concurrent client load. The clean TPC-H import took 29.81 seconds wall time; TPC-DS took 75.83 seconds. A separate client then verified all 32 tables after local staging was gone, including full reads and SHA-256 checks.&lt;/p&gt;

&lt;p&gt;The files are tiny, yet each table’s create, upload, and register step took a median 3.33 seconds for TPC-H and 2.85 seconds for TPC-DS. Those steps account for most elapsed time. This shows a significant per-table fixed cost in this setup; the run does not identify which internal metadata call dominates.&lt;/p&gt;

&lt;p&gt;I then ran the concurrent loader on larger datasets. TPC-H SF 1 loaded 8.66 million rows in 50 seconds; TPC-H SF 10 loaded 86.59 million rows and 3.64 GB of compressed Parquet in 232 seconds. TPC-DS SF 1 loaded 19.56 million rows in 157 seconds; TPC-DS SF 10 loaded 191.50 million rows and 2.77 GB of compressed Parquet across 24 tables in 453 seconds. Each successful run passed independent manifest, footer, and sample Iceberg scan verification. These are single-node wall times, including generation and table registration, not a benchmark result. TPC-H SF 10’s upload phase took about 176 seconds for 3.64 GB, roughly 20.7 MB/s; its streaming/write path needs profiling. A distributed deployment and standard SQL benchmark queries still need separate tests.&lt;/p&gt;

&lt;p&gt;The local Docker volume keeps the Iceberg data when the container stops. To end the example, run &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;docker rm -f crowdb-tpc-demo&lt;/code&gt; and remove &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;/tmp/crowdb-tpc-demo.env&lt;/code&gt;. Remove the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;crowdb-tpc-demo-data&lt;/code&gt; volume only when you deliberately want to discard the imported tables.&lt;/p&gt;
</content></entry><entry><title>Use crowdb-iceberg container with pandas</title><link href="https://buzzcrow.github.io/blog/run-iceberg-in-one-container/"/><id>https://buzzcrow.github.io/blog/run-iceberg-in-one-container/</id><published>2026-09-28T02:30:00+00:00</published><updated>2026-09-29T16:30:00+00:00</updated><summary>A short, runnable path from a Docker container to an Iceberg table and a pandas query.</summary><content type="html">&lt;!-- Publish this revision together with the rebuilt image. The originally published 0.1.0 image does not yet support the append path below. --&gt;

&lt;p&gt;CROWDB’s &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;crowdb-iceberg&lt;/code&gt; container runs an Iceberg REST catalog and its file storage together. This example creates a table of six sample orders, writes the rows through PyIceberg, then reads the stored table into pandas.&lt;/p&gt;

&lt;p&gt;Use disposable data with this &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;0.1.0&lt;/code&gt; evaluation release. The commands below assume a Linux amd64 host and a free local port 80.&lt;/p&gt;

&lt;h2 id=&quot;1-start-the-container&quot;&gt;1. Start the container&lt;/h2&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;docker run &lt;span class=&quot;nt&quot;&gt;-d&lt;/span&gt; &lt;span class=&quot;nt&quot;&gt;--name&lt;/span&gt; crowdb-iceberg &lt;span class=&quot;nt&quot;&gt;-p&lt;/span&gt; 127.0.0.1:80:80 crowdb/crowdb-iceberg:0.1.0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The port mapping lets Python on your machine reach the catalog and file service.&lt;/p&gt;

&lt;h2 id=&quot;2-set-up-the-connection&quot;&gt;2. Set up the connection&lt;/h2&gt;

&lt;p&gt;Install the Python clients and print the connection values from your container:&lt;/p&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;python3 &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; venv .venv
&lt;span class=&quot;nb&quot;&gt;.&lt;/span&gt; .venv/bin/activate
pip &lt;span class=&quot;nb&quot;&gt;install&lt;/span&gt; &lt;span class=&quot;s1&quot;&gt;&apos;pyiceberg[pyarrow]==0.11.1&apos;&lt;/span&gt; pandas
docker &lt;span class=&quot;nb&quot;&gt;exec &lt;/span&gt;crowdb-iceberg crowdb-monitor credentials show &lt;span class=&quot;nt&quot;&gt;--format&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;env&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Export the two Iceberg values in the same shell, replacing the examples below with the values printed by your container. Keep the token private.&lt;/p&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;ICEBERG_URI&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;http://localhost&apos;&lt;/span&gt;
&lt;span class=&quot;nb&quot;&gt;export &lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;ICEBERG_TOKEN&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;&amp;lt;your container token&amp;gt;&apos;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Start an &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;orders.py&lt;/code&gt; file with the connection:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;

&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pandas&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pd&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pyarrow&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pa&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;from&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pyiceberg.catalog&lt;/span&gt; &lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;load_catalog&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;catalog&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;load_catalog&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
    &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;crowdb&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;nb&quot;&gt;type&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;rest&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;uri&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;environ&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;ICEBERG_URI&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;token&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;environ&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;ICEBERG_TOKEN&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;3-write-a-table-and-its-data&quot;&gt;3. Write a table and its data&lt;/h2&gt;

&lt;p&gt;Add this block to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;orders.py&lt;/code&gt;. The sample rows are created locally; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;table.append&lt;/code&gt; writes them to the container’s Iceberg storage.&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;orders&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pd&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nc&quot;&gt;DataFrame&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
    &lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;Beijing&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;paid&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;120&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;Shanghai&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;paid&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;80&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;Beijing&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;cancelled&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;200&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;4&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;Shanghai&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;paid&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;60&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;5&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;Shenzhen&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;paid&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;50&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;6&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;Beijing&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;paid&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;30&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;
    &lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;columns&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;order_id&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;city&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;status&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;amount_usd&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;],&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;arrow_orders&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pa&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Table&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;from_pandas&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;orders&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;preserve_index&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;False&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;catalog&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;create_namespace_if_not_exists&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;pandas_demo&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;table&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;catalog&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;create_table&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;pandas_demo.orders&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;schema&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;arrow_orders&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;schema&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;table&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;append&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;arrow_orders&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This creates a new &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;pandas_demo.orders&lt;/code&gt; table. Use a different table name if you run the script again.&lt;/p&gt;

&lt;h2 id=&quot;4-query-the-stored-table-with-pandas&quot;&gt;4. Query the stored table with pandas&lt;/h2&gt;

&lt;p&gt;Add the final block to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;orders.py&lt;/code&gt;, then run &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python orders.py&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;saved_orders&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;catalog&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;load_table&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;pandas_demo.orders&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;scan&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;().&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;to_pandas&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;result&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;saved_orders&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;saved_orders&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;status&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;paid&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
    &lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;groupby&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;city&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;as_index&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;False&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;agg&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;orders&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;order_id&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;count&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;revenue_usd&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;amount_usd&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;sum&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;
    &lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;sort_values&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;revenue_usd&lt;/span&gt;&lt;span class=&quot;sh&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;ascending&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;False&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;nf&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;result&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;to_string&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;index&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;False&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;python orders.py
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The query counts paid orders and sums their revenue for each city:&lt;/p&gt;

&lt;div class=&quot;language-text highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;    city  orders  revenue_usd
 Beijing       2          150
Shanghai       2          140
Shenzhen       1           50
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The cancelled order is excluded. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;saved_orders&lt;/code&gt; comes from a fresh Iceberg table scan, so this query uses the data written to the container.&lt;/p&gt;

&lt;p&gt;For container configuration and other client operations, see the &lt;a href=&quot;https://crowdb.dev/docs/quickstart/&quot;&gt;quick start&lt;/a&gt; and &lt;a href=&quot;https://crowdb.dev/docs/manual/iceberg/&quot;&gt;Iceberg manual&lt;/a&gt;.&lt;/p&gt;
</content></entry><entry><title>Why we’re building CROWDB</title><link href="https://buzzcrow.github.io/blog/why-we-are-building-crowdb/"/><id>https://buzzcrow.github.io/blog/why-we-are-building-crowdb/</id><published>2026-09-21T02:30:00+00:00</published><updated>2026-09-29T02:00:00+00:00</updated><summary>Storage bottlenecks move. System boundaries tend to stay. CROWDB is an attempt to own enough of the data path to change both.</summary><content type="html">&lt;p&gt;For more than a decade, I have designed, built, and debugged storage systems as their bottlenecks moved from disks to CPUs, networks, and data movement. The most useful lessons came from the work in between: revisiting an assumption, tracing a slow request, and seeing what a real workload exposed that a clean diagram did not.&lt;/p&gt;

&lt;p&gt;That experience is why I’m building CROWDB. I want the storage system to own enough of the data path that we can change it when the workload changes—not just tune the parts we happen to control.&lt;/p&gt;

&lt;h2 id=&quot;storage-has-a-new-job&quot;&gt;Storage has a new job&lt;/h2&gt;

&lt;p&gt;Consider a training job that reads data from an Iceberg table. The table engine resolves metadata, reads files, and hands data to a library that prepares batches for the GPU. Somewhere underneath, storage places and protects the bytes. Several components may cache, retry, or copy the same data along the way.&lt;/p&gt;

&lt;p&gt;None of those steps is automatically wrong. Separate systems give us useful interfaces and independent ways to evolve. The difficult part is deciding who owns a problem that crosses those interfaces. A table engine cannot change how storage places a file. A storage service usually cannot see the batch the application is trying to assemble.&lt;/p&gt;

&lt;p&gt;When I follow a slow request, those boundaries are often more interesting than the individual algorithms. Which component owns this buffer? Why is this metadata fetched again? What happens if the writer retries after only part of the operation became durable? A fast device does not answer those questions.&lt;/p&gt;

&lt;p&gt;S3 remains a useful interface. I do not think every workload should be forced to use it as an intermediate representation, though. Tables have commits and snapshots. Datasets have samples and batches. These concepts deserve a place in the storage design, not just conventions layered over object names.&lt;/p&gt;

&lt;h2 id=&quot;what-crowdb-is&quot;&gt;What CROWDB is&lt;/h2&gt;

&lt;p&gt;CROWDB is a distributed storage platform. S3 and Iceberg have their own access models over a shared storage core. Dataset is a third model we are designing; it is not implemented yet.&lt;/p&gt;

&lt;p&gt;An S3 client works with objects. An Iceberg client works with a catalog, tables, snapshots, and immutable files. The two models reuse the same underlying machinery for distributed state and protected chunks. Iceberg is not implemented by asking users to assemble a catalog on top of a separate S3 service.&lt;/p&gt;

&lt;p&gt;This distinction is easy to lose in a diagram. “Shared storage” does &lt;strong&gt;not&lt;/strong&gt; mean that uploading an S3 object registers an Iceberg table. The access models still have separate semantics. What they share is the infrastructure underneath.&lt;/p&gt;

&lt;p&gt;The first useful question is therefore practical: can a table client talk to the catalog and store its data without another storage service to configure? The native Iceberg path is designed to make that possible. The answer still needs to be tested against each client workflow.&lt;/p&gt;

&lt;h2 id=&quot;why-build-the-whole-path&quot;&gt;Why build the whole path?&lt;/h2&gt;

&lt;p&gt;There are good storage engines, object stores, and table formats already. Reusing them is often the right decision. Building another core means taking responsibility for failures that someone else has already spent years finding.&lt;/p&gt;

&lt;p&gt;I am accepting that cost because the boundaries are part of what I want to change. If a request crosses an independently owned metadata service, an object gateway, and a separate storage engine, a change to placement or buffer ownership can become a negotiation between systems. We want those contracts to be explicit within CROWDB.&lt;/p&gt;

&lt;p&gt;Multi-Paxos, write-ahead logs, B+trees, and erasure coding are not the novelty. Putting them in one repository is not an advantage by itself. The work is making their contracts agree: when a write is durable, who may change its placement, how a buffer stays bounded, and what recovery must finish after a process dies.&lt;/p&gt;

&lt;p&gt;Owning those decisions should give us room to improve the path. It does not prove that CROWDB is faster. That needs measurements with a stated workload, hardware, concurrency, and baseline. I would rather publish those results separately than borrow credibility from an architecture drawing.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Own the path that determines the system’s limits.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;three-layers-one-system&quot;&gt;Three layers, one system&lt;/h2&gt;

&lt;p&gt;The architecture has three layers. Reading from the application downward helps explain their jobs, but this is a dependency map—not a claim that every byte travels through every box.&lt;/p&gt;

&lt;figure class=&quot;architecture-figure&quot; aria-label=&quot;CROWDB architecture: access models, chunk layer, and reusable KV&quot;&gt;
  &lt;div class=&quot;figure-header&quot;&gt;&lt;span&gt;FIG. 01 / THREE LAYERS&lt;/span&gt;&lt;span class=&quot;figure-legend&quot;&gt;&lt;span&gt;&lt;i&gt;&lt;/i&gt;implemented&lt;/span&gt;&lt;span&gt;&lt;i class=&quot;dashed&quot;&gt;&lt;/i&gt;planned&lt;/span&gt;&lt;/span&gt;&lt;/div&gt;
  &lt;div class=&quot;access-cards&quot;&gt;
    &lt;div class=&quot;access-card&quot;&gt;&lt;b&gt;S3&lt;/b&gt;&lt;span&gt;HTTP objects&lt;/span&gt;&lt;span class=&quot;tiny-status&quot;&gt;IMPLEMENTED&lt;/span&gt;&lt;/div&gt;
    &lt;div class=&quot;access-card&quot;&gt;&lt;b&gt;Iceberg&lt;/b&gt;&lt;span&gt;Catalog + FileIO&lt;/span&gt;&lt;span class=&quot;tiny-status&quot;&gt;IMPLEMENTED&lt;/span&gt;&lt;/div&gt;
    &lt;div class=&quot;access-card planned&quot;&gt;&lt;b&gt;Dataset&lt;/b&gt;&lt;span&gt;HTTP + native client&lt;/span&gt;&lt;span class=&quot;tiny-status&quot;&gt;IN DESIGN&lt;/span&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class=&quot;figure-connector&quot; aria-hidden=&quot;true&quot;&gt;↓&lt;/div&gt;
  &lt;div class=&quot;layer-box dark&quot;&gt;&lt;div class=&quot;layer-title&quot;&gt;&lt;span&gt;02&lt;/span&gt;Chunk&lt;/div&gt;&lt;div class=&quot;layer-main&quot;&gt;&lt;b&gt;One protected storage layer&lt;/b&gt;&lt;p&gt;Chunk Stream · Chunk-KV&lt;br /&gt;Placement, bounded streaming, protection and repair&lt;/p&gt;&lt;div class=&quot;path-line&quot;&gt;chunk client → Chunk I/O → ChunkDB → DiskIO → DiskDB&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
  &lt;div class=&quot;figure-connector&quot; aria-hidden=&quot;true&quot;&gt;↕&lt;/div&gt;
  &lt;div class=&quot;layer-box&quot;&gt;&lt;div class=&quot;layer-title&quot;&gt;&lt;span&gt;01&lt;/span&gt;Reusable KV&lt;/div&gt;&lt;div class=&quot;layer-main&quot;&gt;&lt;b&gt;Replicated metadata and durable state&lt;/b&gt;&lt;p&gt;crowdb-kv · Multi-Paxos · WAL · crowdb-tree · RPC&lt;/p&gt;&lt;/div&gt;&lt;/div&gt;
  &lt;div class=&quot;planned-path&quot;&gt;&lt;strong&gt;PLANNED PATH&lt;/strong&gt; &amp;nbsp; DiskIO buffers → RDMA / GDS → GPU memory. This is a design direction, not a capability of the preview.&lt;/div&gt;
  &lt;figcaption&gt;The KV layer supplies distributed state; the diagram does not imply that every payload is stored in KV. S3 and Iceberg use the shared core without translating through one another.&lt;/figcaption&gt;
&lt;/figure&gt;

&lt;h3 id=&quot;access-models-keep-their-meaning&quot;&gt;Access models keep their meaning&lt;/h3&gt;

&lt;p&gt;The Access layer implements the interfaces applications use. S3 handles HTTP object operations. Iceberg supplies a native catalog and FileIO path. Both have working implementations in the current repository. Dataset remains in design, including the possibility of native access that does not pass through an HTTP server.&lt;/p&gt;

&lt;p&gt;The Iceberg catalog and FileIO belong to one access model. General S3 object access is a separate model. Keeping those roles clear is more useful than calling every path “compatible.”&lt;/p&gt;

&lt;h3 id=&quot;chunks-own-the-storage-work&quot;&gt;Chunks own the storage work&lt;/h3&gt;

&lt;p&gt;The Chunk layer is responsible for placement, protection, streaming, and repair. Chunk Stream provides durable ordered append; Chunk-KV provides a range-partitioned structure. The underlying path continues through Chunk I/O, ChunkDB, DiskIO, and DiskDB.&lt;/p&gt;

&lt;p&gt;These names will need their own articles. For this introduction, the important point is ownership: access models can reuse the same storage mechanisms rather than grow their own placement and recovery systems.&lt;/p&gt;

&lt;h3 id=&quot;distributed-state-is-a-reusable-foundation&quot;&gt;Distributed state is a reusable foundation&lt;/h3&gt;

&lt;p&gt;Underneath, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;crowdb-kv&lt;/code&gt; supplies replicated state using parallel Multi-Paxos slots and write-ahead logging. It is a reusable distributed layer, not the product-level data model.&lt;/p&gt;

&lt;p&gt;This separation matters to me. Owning the stack should not require turning it into a single inseparable component. The layers need narrow enough contracts that we can reason about them, test them, and use them independently where that makes sense.&lt;/p&gt;

&lt;h2 id=&quot;built-for-the-next-data-path&quot;&gt;Built for the next data path&lt;/h2&gt;

&lt;p&gt;The GPU path is an example of why I want this control. Today, CROWDB uses ordinary CPU streaming. Dataset, topology-aware native access, and direct GPU delivery remain planned work.&lt;/p&gt;

&lt;p&gt;RDMA or GPUDirect Storage would not become useful merely because we added another endpoint. We would need to reason about buffer ownership, placement, protection, and the lifetime of the data being transferred. Those decisions cross several layers of the system.&lt;/p&gt;

&lt;p&gt;The aim is to leave room for that work without changing what an object or a table means. There is no direct-GPU benchmark to show here, and no such capability to enable in the preview. That boundary belongs next to the design, not in a footnote.&lt;/p&gt;

&lt;h2 id=&quot;where-the-project-stands&quot;&gt;Where the project stands&lt;/h2&gt;

&lt;div class=&quot;callout&quot;&gt;
&lt;strong&gt;Development status · updated September 29, 2026&lt;/strong&gt;
&lt;p&gt;S3 and native Iceberg have working implementations. Dataset and direct GPU delivery are still in design. Production use and on-disk upgrade compatibility are not supported.&lt;/p&gt;
&lt;/div&gt;

&lt;p&gt;The repository includes code, tests, and design documents; the product site hosts the user manual. That is where I want this argument to be judged. A useful next step is to inspect the Iceberg guide and tell us which assumption breaks under your workload.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://crowdb.dev/docs/quickstart/&quot;&gt;Start with the Iceberg evaluation guide →&lt;/a&gt;&lt;/p&gt;

&lt;div class=&quot;source-note&quot;&gt;Technical references: &lt;a href=&quot;https://github.com/buzzcrow/crowdb&quot;&gt;project README and status&lt;/a&gt;, &lt;a href=&quot;https://github.com/buzzcrow/crowdb/tree/main/doc/design&quot;&gt;design documents&lt;/a&gt;, and &lt;a href=&quot;https://crowdb.dev/docs/quickstart/&quot;&gt;Iceberg quick start&lt;/a&gt;. This article was first published on September 21 and revised on September 29, 2026 to reflect the current Iceberg implementation and development limits.&lt;/div&gt;
</content></entry>
</feed>