17 September 2026 · 7 min
Cold starts and long runs
Spark, Fabric, BigQuery slots, Oracle parallel query - they all charge you for waking up. A pipeline that cold-starts every five minutes for a 90-second transform will spend more time booting than working. We keep a small warm pool only for the path the dashboard actually hits. Everything else is batch, with a published SLA.
Long-running jobs need a kill switch and a checkpoint. A 6-hour load that fails at 5 hours 50 and starts from zero is how weekends disappear. Micro-batches with a watermark beat one heroic run.
What 'near live' should mean
We agree the grain and the lag in writing: payments within 15 minutes, headcount overnight, forecast weekly. Then we build only that. A private-sector ops team wanted 'real time' and was paying for streaming on data that changed twice a day. We moved them to a 15-minute increment plus a cache. The dashboard felt live. The bill did not.
Related service: Data insights & strategy