Archives
Browse all archived articles.
2026 31
September 2
-
Data This Week #31
5 min readPandas vs. single-machine engines, Delta Lake 4.3 selective data replacement, PlanetScale's Neki, PyIceberg with Apache Polaris, Shopify's gisting, Sony LIV's real-time analytics, and a community debate on dbt development standards.
-
Data This Week #30
7 min readAmazon acquires DuckDB Labs, AngelList's self-updating Markdown semantic layer, managing AI coding costs at Databricks, federated MCP access for AI agents, Glue 6.0 real-time mode, Polars 2.0 RC, and community advice on extracting from legacy on-prem databases.
August 4
-
Data This Week #29
5 min readDuckDB v2.0's new PEG-based SQL parser, Cassandra's ACCORD consensus protocol for ACID transactions, migrating multilingual full-text search to PostgreSQL, stream deduplication techniques, the hidden joins inside Snowflake MERGE, and the community debate on enterprise data catalog alternatives.
-
Data This Week #28
5 min readInstacart's Postgres-native search replacing Elasticsearch and FAISS, Netflix migrating batch workloads to Kueue on Kubernetes, AWS Dogwood for AI agent governance, distributed DuckDB with Quack, Debezium for CDC-driven event architecture, and the community debate on dbt Cloud pricing.
-
Data This Week #27
6 min readSafely migrating Iceberg catalogs without moving data, DuckDB's MySQL engine stress-tested at 500 GB, the end of the Hive Metastore era, hands-on S3 Tables evaluation, why dbt tests don't equal data trust, and an ontology+LLM approach to legacy data modernization.
-
Data This Week #26
4 min readAmazon MSK delivers Kafka data directly to Iceberg streaming tables, Iceberg v3 introduces native row-level lineage for CDC, Spotify builds a custom indexing layer for online point queries on the data lake, and a Rust terminal database GUI called Rainfrog.
July 4
-
Data This Week #25
6 min readS3 Tables compaction myths debunked, DuckDB's vectorized execution internals, Snowflake's managed Iceberg at scale, NOT EXISTS rewritten as anti-joins, partition affinity over Redis, AWS Glue view automation, and Supabase Pipelines in public alpha.
-
Data This Week #24
5 min readDatabricks bakes native AI and semantic primitives into Spark 4.2, Netflix rebuilds LLM serving with vLLM on Triton, multi-cloud lakehouse architecture on AWS for agentic AI, Kubernetes internals deep dive, and Debezium's log-based CDC architecture.
-
Data This Week #23
5 min readApache OSSIE enters ASF incubation to standardize semantic layers, HubSpot scales to 20B vectors with Qdrant, cloud-native financial search with Iceberg and Turbopuffer, Lakekeeper's Generic Table API for multi-format lakehouses, and versioning Power BI with Git.
-
Data This Week #22
5 min readAWS S3 Annotations for business context, Cloudflare's Town Lake lakehouse and Skipper AI agent, Databricks LTAP rethinking database storage, Iceberg native views in Hive, Snowflake pipeline scaling pitfalls, and CRED's zero-data-loss RDS Blue/Green deployments at scale.
June 4
-
Data This Week #21
4 min readPostgres 19 pg_plan_advice for query plan control, Flink's Hadoop-free native S3 filesystem, incremental model pitfalls in dbt, Trino's summer SQL standard upgrades, PgBouncer's pooling mechanics, Razorpay's CDP architecture, and pgEdge ColdFront for Iceberg tiering.
-
Data This Week #20
5 min readKafka's nine-layer architecture breakdown, self-healing pipeline barriers, ClickHouse full-text search on object storage, Google Cloud Next '26 data infrastructure, watsonx.data semantic layer, Neo4j on Databricks, DuckDB internals, and handling messy Excel ETL.
-
Data This Week #19
5 min readSlack's SSH-to-REST EMR migration, clusterless Iceberg Lakehouse with DuckDB, Iceberg v4 metadata proposals, Spark 4.0 on EMR GA, Apache Gravitino unified catalog, and Databricks Omnigent for AI agent orchestration.
-
Data This Week #18
4 min readSupabase Multigres for Postgres scaling, Netflix's dynamic Cassandra partition splitting, dbt Core v2 on Rust, SQL/PGQ graph queries in Postgres 19, and IAM fundamentals for data engineers.
May 5
-
Data This Week #17
4 min readPulsar 5.0 Scalable Topics, Iceberg multi-engine query routing, Netflix's high-throughput graph abstraction, Iceberg partition evolution, AWS Glue automation, and community thoughts on Databricks' BI migration tool.
-
Data This Week #16
5 min readDatabricks as a unified platform, PostgreSQL pluggable storage engines, Iceberg 1.11.0 highlights, database egress costs, JDBC query caching, and dbt column-level lineage.
-
Data This Week #15
6 min readSpark Declarative Pipelines for financial lakehouses, ten AWS Glue & Iceberg fixes, MOR as an architectural shift, DuckDB's Quack protocol, SQL fraud patterns, Kafka checkpoint patterns, and the LLM-for-validation debate.
-
Data This Week #14
6 min readFlink CDC streaming ELT from MySQL to Kafka, the LLM engineer's stack map, Ursa's diskless Kafka fork, Iceberg write mechanics, Instacart's billion-product search, Jikkou 1.0, and the AI knowledge-base debate.
-
Data This Week #13
5 min readSpark memory tuning, row-level validation tiers, Postgres RLS pitfalls, Stripe's sharding at 5M QPS, Aurora DSQL vs Postgres, Velero joins CNCF, and SQLGlot 5x faster with mypyc.
April 4
-
Data This Week #12
5 min readCold Postgres data to S3 lakehouse, Databricks Lakeflow Designer, vector databases & HNSW indexing, Salesforce migration best practices, SwiftLake for Iceberg, and data observability lessons.
-
Data This Week #11
5 min readIceberg cross-account migrations, DuckLake 1.0 metadata, IaC for data engineers, Redshift Iceberg writes, agent-data patterns, LARQL for LLM graph queries, and Dagster pricing debate.
-
Data This Week #10
5 min readData product lifecycle, semantic context layer for LLM agents, Netflix's Druid interval caching, Ursa Kafka storage engine, Iceberg v3 VARIANT type, and Ministack vs LocalStack.
-
Data This Week #9
5 min readDuckLake's 926x Iceberg speedup, Expedia's Trino Gateway for workload routing, Ontul unified SQL engine, PostgreSQL memory myths, and a 6-tier FFLIIP streaming lakehouse deep-dive.
March 4
-
Data This Week #8
4 min readPydantic for schema contracts, Databricks Vector Search pitfalls, stateless Kafka broker Tansu, Capital One's GenAI agent, RAG as a DE problem, and testing culture in data teams.
-
Data This Week #7
4 min readNetflix's RDS-to-Aurora PostgreSQL migration, DuckDB cost optimization, real-time dashboards with LISTEN/NOTIFY, Airflow on Minikube, and the dbt vs. SQLMesh debate in 2026.
-
Data This Week #6
4 min readXiaomi's unified lakehouse with Doris & Paimon, Top-K in Postgres, dbt run monitoring, PostgreSQL internals, Netflix's DataJunction semantic layer, and schema evolution debates.
-
Data This Week #5
4 min readSpark DAG compilation deep dive, query federation with StarRocks, Pinterest's CDC migration, CyberArk AI with Iceberg, Databricks Zerobus Ingest, and data quality tooling debates.
February 4
-
Data This Week #4
5 min readHow OpenAI scales PostgreSQL for ChatGPT, Dropbox's enterprise RAG, 3x faster Spark on Iceberg, dbt with DuckDB, local AWS Lakehouse setups, and new tool Alibaba ZVec.
-
Data This Week #3
3 min readBigQuery cost optimization, Apache Iceberg updates, MinIO alternatives, AWS SageMaker governance, and new tools like Nao — curated for data engineers.
-
Data This Week #2
4 min readRisingWave HTTP streaming to Iceberg, CedarDB string compression, Alibaba open-sources AliSQL (MySQL + DuckDB), Databricks Lakebase GA, and AI-powered data quality monitoring.
-
Data This Week #1
6 min readPostgreSQL dominance in 2025, Arrow-based database connectivity, Uber's petabyte-scale replication, Netflix AI graph search, and new tools OpenEverest and Pandas 3.0.