Tag: Data Engineering
All the articles with the tag "Data Engineering".
- 30 MIN READ•Jul 6, 2026
The State of Apache Iceberg v4 in July 2026: What the Dev List Tells Us About the Format's Next Chapter
What the Iceberg v4 dev list tells us about adaptive metadata trees, single-file commits, column updates, and the format's next chapter in mid-2026.
Apache Icebergdata engineeringlakehouse architecture - 29 MIN READ•Jul 6, 2026
The State of Agentic AI Standards in 2026: MCP, A2A, WebMCP, OSI, and the Protocol Stack Taking Shape
The agentic AI protocol stack is solidifying in 2026 — MCP for tools, A2A for agents, WebMCP for the web, OSI for semantics, payments, identity, and security.
Apache Icebergdata engineeringlakehouse architecture - 28 MIN READ•Jul 6, 2026
The State of Apache Arrow in 2026: Ten Years In, the Invisible Standard Is Everywhere
Apache Arrow at 10 — ADBC, Flight SQL, nanoarrow, the AI reinterpretation, and how an in-memory standard eliminated the copy tax across the data stack.
Apache Arrowdata engineeringlakehouse architecture - 28 MIN READ•Jul 6, 2026
The State of Apache Parquet in 2026: The Quiet Format Enters Its Loudest Decade
Apache Parquet in 2026 — variant types, geospatial, ALP encoding, footer redesign, the versioning debate, and how the decade-old format is renovating for AI workloads.
Apache Parquetdata engineeringlakehouse architecture - 29 MIN READ•Jul 6, 2026
The State of Apache Polaris in July 2026: From Incubating Catalog to the Governance Layer of the Open Lakehouse
Apache Polaris as a TLP — federation, credential vending, semantic layers, lineage, and how the open catalog became the governance plane of the multi-engine lakehouse.
Apache Polarisdata engineeringlakehouse architecture - 30 MIN READ•Jul 6, 2026
The State of Streaming to Apache Iceberg in July 2026: Every Path, Its Latency, and What to Do When Seconds Are Not Fast Enough
Every path for streaming data into Iceberg in 2026 — Flink, Spark, Kafka Connect, broker-native, managed pipelines — with honest latency numbers and sub-second hybrid architectures.
Apache Icebergstreamingdata engineering - 26 MIN READ•Jul 6, 2026
Deterministic Data Engineering With AI Harnesses: Using Claude Code, Codex, Antigravity, and OpenCode for Data Work You Can Actually Trust
How to use AI agent harnesses for data engineering without losing determinism, reproducibility, and trust in your data pipelines and analytics.
data engineeringAI agentsClaude Code - 14 MIN READ•Jun 8, 2026
Apache Iceberg v4 Roadmap: Adaptive Metadata Trees, Single-File Commits, and the Delta Convergence
A deep technical breakdown of Apache Iceberg v4's proposed architecture: adaptive metadata trees, one-file commits, relative paths, column families, and what the Delta 5.0 convergence actually means for your data platform.
apache-icebergopen-table-formatsdata-engineering - 16 MIN READ•May 31, 2026
Data Platform Native AI Agent Tooling in 2026
A comprehensive comparison of AI agent tooling across Dremio, Snowflake, Databricks, Microsoft Fabric, AWS, Google Cloud, ClickHouse, VeloDB, SpiceAI, Bauplan, and Qlik.
AI AgentsData PlatformsMCP - 24 MIN READ•May 23, 2026
Single-Node Data Engineering: DuckDB, DataFusion, Polars, and LakeSail
Optimize single-node data engineering with DuckDB, DataFusion, Polars, and LakeSail. Compare architectures and learn when to transition to Dremio MPP.
DuckDBApache ArrowDataFusion - 27 MIN READ•May 22, 2026
Apache Iceberg SCD Type 2 and CDC Patterns: Building Historical Lakehouse Tables
A deep dive into implementing Slowly Changing Dimension Type 2 (SCD Type 2) patterns and Change Data Capture (CDC) pipelines on Apache Iceberg, using PySpark and Dremio.
apache icebergcdcscd type 2 - 24 MIN READ•May 22, 2026
Apache Iceberg Catalogs Explained: REST, Glue, Hive Metastore, Polaris, Nessie, and Snowflake
A deep dive into Apache Iceberg catalog architecture, comparing REST catalogs, AWS Glue, Project Nessie, Polaris, and Snowflake. Learn catalog role, credential vending, and cross-engine configurations.
apache icebergcatalogsNessie