Posts

Life stuff. What I'm up to, or thinking about. Or just thing's that I want to write down so I don't forget.

Spark Fundamentals: Partitioning

Two kinds of partitioning — logical (in-memory) and physical (on-disk) — and the rules of thumb for getting each one right: sizing memory partitions for parallelism, and sizing/ordering storage partitions for fast scans.

On the list

Topics I'm planning to write about next:

  • Data Platform Orchestration: Databricks and Airflow
  • Dead Letter Queues
  • LLM-assisted Data Platforms development
  • Logging and Monitoring
  • Microsoft Fabric Workspace Strategies
  • Schema Change Management Strategies
  • Setting up CDC against SQL Databases
  • Should our dataplatform run with a GUI? Or a Library?
  • Turtle Enclosure Stream Setup

Junusais quoi

(Claude-generated placeholder.) This site — an Astro static build with a Markdown blog and portfolio via content collections, plus a turtle-cam page streaming a live Tapo camera over HLS. The camera feed runs through MediaMTX (RTSP→HLS) and a Cloudflare Tunnel on a separate always-on machine, with no exposed home IP or port forwarding.

Cosmos Flink streaming lab

(Claude-generated placeholder.) Azure Cosmos DB's change feed piped through Kafka into a Flink SQL job doing windowed per-customer aggregation. The point wasn't just getting it running — it's a working-through of schema evolution under Schema Registry compatibility constraints, retention as a recovery-time budget, and why teams replay from a durable source instead of leaning on a DLQ.

Notion

(Claude-generated placeholder.) Syncs a Goodreads library export into a Notion database — parses the CSV, builds an ISBN13 lookup (falling back to title+author when Goodreads leaves it blank), and upserts each book against the existing Notion pages, paginating through the database and backfilling missing ISBNs along the way.