TLDR Data 2026-10-01
Claude Kills the Claudisms βοΈ, X-ray Your Parquet π©», DuckDB 1.5.6 Lands π¦
Row-level security performance in PostgreSQL, measured (9 minute read)
PostgreSQL row-level security can be effectively free when the policy is a simple indexed tenant check, but poorly structured policies can turn sub-millisecond queries into tens of milliseconds or even seconds. The main traps are per-row membership lookups, VOLATILE helper functions, and non-leakproof query functions that stop PostgreSQL from using the full index.
Rebuilding Data Engineering with Harness Engineering: A New Paradigm for the Agent Era (18 minute read)
Data platforms are shifting from human-operated tooling to agent-ready infrastructure, where the core challenge is not generating SQL or pipelines but safely turning AI-generated work into production outcomes. The proposed stack centers on five layers: Intent, Control, Semantic, Harness, and Runtime. For data teams, the implication is a move from building scripts and DAGs to designing reusable Skills, Policies, and context-rich execution loops that make automation trustworthy.
Improving Cost Efficiency of Data Streaming Pipelines (12 minute read)
Streaming platforms get expensive for structural reasons, not data volume: request counts, extra copies of each message, serialization, and cross-AZ traffic often dominate the bill. Practical wins include batching producers and consumers, eliminating intermediate topics, using columnar payloads and compression, paying for low latency only where workloads truly need it, designing state deliberately, and keeping cost attribution visible per pipeline.
Can Your LLM Vibe Code Your Semantic Layer? (6 minute read)
A strong semantic layer can dramatically improve LLM analytics accuracy, with the best model-built layer reaching 91% versus 39% with no layer. The main failure mode was not metric arithmetic but hidden business conventions like which accounts count as customers, showing that semantic layers still need explicit human-defined rules and validation.
RIP, vector database (4 minute read)
turbopuffer v3 stops using the ANN vector index as the primary storage layout. That should reduce duplicated data and expensive rewrites, while letting full-text, filtering, aggregations, and vector search each use storage and block sizes better suited to their workload.
Handling hot shards (6 minute read)
Sharding by tenant ID works until a large tenant makes one shard run hot, as Slack learned when enterprise workspaces overwhelmed workspace-keyed shards. The fixes range from vertical scaling to isolating whale tenants on dedicated shards and ultimately resharding tables by the column queries actually touch. Choosing keys per table rather than globally keeps common access paths single-shard while spreading load.
The Five Camps of Data Modeling (and Which One You're Stuck In) (6 minute read)
Data modeling grew up in five separate camps: relational, analytics, application, ML/AI, and knowledge/ontology, each solving different problems with real blind spots. One JSON product catalog can satisfy an application while breaking analytics, ML, and governance, producing brittle pipelines and lost dashboard trust. Modern stacks span all five camps at once, so shared literacy across them matters more than any single modeling tradition.
Precisely Now: Where Data Meets AI β Oct 8, Free (Sponsor)
Bad data is the #1 reason AI agents fail before they ship. At Precisely Now, see our new AI builder experience live: ready-made agents, skills, and prompts for teams who build in the real world.
Register Now.Parquet X-ray (GitHub Repo)
Parquet X-ray is a browser-based visual inspector for understanding how a Parquet file is physically laid out. You can open a local file or remote URL and explore row groups, column chunks, pages, indexes, bloom filters, schema, min/max statistics, compression efficiency, and byte-level layout without downloading the whole file. You can test it out via Hugging Face.
Announcing DuckDB 1.5.6 (3 minute read)
DuckDB 1.5.6 is a maintenance release with correctness, crash, security, and performance fixes. DuckDB 2.0 is in development, with an internal Windows TPC-H test running over 6x faster than 1.5.6, though DuckDB says that speedup will not generalize to every workload.
Splink 5: Probabilistic record linkage at billion-row scale (7 minute read)
Splink, an open-source library for record linkage and deduplication, scales probabilistic record linkage to billion-row jobs: a prediction job ran 10 billion comparisons in 8.5 minutes on a 192-vCPU instance using the upcoming DuckDB 2.0 engine. The tool supports chunked sampling to speed up model training or distribute processing and incremental prediction APIs to link new records without full re-runs.
A Hint of Dependence (14 minute read)
Query hints solve bad plans by hard-coding execution choices into SQL, which creates brittle coupling to indexes and data shape. A better approach is to keep plan control outside the query and improve the planner's information with tools like plan management and CREATE STATISTICS, so it can adapt as the data changes.
Anthropic says it fixed Claude's writing. I ran the evals to check (12 minute read)
Anthropic's writing changes in Opus 5.5 appear to work partly: em dashes fell by 99.6% and other recognizable βClaudismsβ fell about 50% versus Opus 5. The bigger lesson is methodological: style evals work better when you annotate exact problematic spans and judge recurring rhetorical moves, not just banned phrases or vague overall scores.
Curated deep dives, tools and trends in big data, data science and data engineering π
Join 590,000 readers for
one daily email