Book a 30-minute audit

Four days out · the site, empty

4 days before any of this opened.

The architecture was locked with four days left. Eleven platforms to read, eighty-nine numbers to define, and a festival that happens exactly once.

The brief · 89 numbers

11 platforms. 89 numbers. 1 deadline.

Each platform with its own API, its own auth model, and its own idea of when a day starts. Reconciling them into one comparable day is most of the work.

12:00 · the deadline

And every morning, all of it on one screen by noon.

Collectors fire at 12:05. Manual entries close at 12:45. A watchdog checks that every source reported and emails a pass or a fail.

Day 1 · gates open

Day 1. The system was 3 days old.

125k people through the gates on the first day alone. Nothing had been tested at that volume, because nothing could be.

Day 1 · in production

Then it stopped being a plan.

Roughly half the final system was built after the gates opened — in production, while it was already reporting.

Day 2 · 10:20 · presale opens

Day 2: the next edition went on sale, mid-festival.

Fifteen individually tagged launch links, live with the site already full. A nine-day-old system attributing sales by channel while the campaign ran.

Four days · the whole site

500k people, present on the festival grounds.

466,911 individual values across 17 days, 21 workflows, and sixteen failures written down rather than buried.

Measured while it happened.

Case study · UNTOLD, Cluj-Napoca

The plan was locked
four days before
the gates opened.

UNTOLD, in Cluj-Napoca — one of the top 3 festivals in the world, and 500k people over 4 days. Eighty-nine numbers, across eleven platforms, on a screen by noon every single morning — with the next edition's ticket launch buried in the middle of it.

30 Jul – 15 Aug 2026 17 days of collection 466,911 datapoints one person, no data team

The film above is built from UNTOLD's own photographs, converted to black and white. One frame — the production desk — is generated, because nobody photographs that room.

0
Individual values collected
0
Workflows, 539 steps
0
Metrics defined, 80 live
0
External platforms
0
Days from locked plan to gates
0
Failures written down, not buried

The brief

Every morning,
before noon.

The marketing team needed to know how the festival was performing across every channel it touched, and to be able to put that in front of leadership by midday. 89 metrics, refreshed daily. 11 external platforms, each with its own API, auth model and quirks. 4 consecutive festival days plus rehearsals. A hard noon deadline. And no tolerance for a wrong number, because the output fed leadership decks.

The constraints were sharper than the brief. No data engineering team — one person working with AI assistance, plus a marketing team entering a handful of values by hand. No time: the architecture was locked four days before the gates opened. And no second chance, because a festival happens once and most social APIs will not give you yesterday's story counts after the fact.

  1. A “day” is not a day

    Every platform defines its reporting window differently, in a different timezone, with different cut-offs. Reconciling them into one comparable business day is most of the work.

  2. Data is not final when you ask for it

    Web analytics backfill for 24–48 hours. Social posts keep accumulating views for days. A number read at noon is not the number that will be true at midnight.

  3. Platforms delete metrics

    Quietly, on a published schedule almost nobody reads. There is no error and no email. The field simply stops appearing.

  4. Platforms lie by omission

    They will show a number in their own dashboard and expose it through no API, at any price. Four metrics here were impossible from day one.

  5. Volume is not constant

    Off-season these accounts post 5–14 times a day. During the festival, 250. Every limit and cap set in a quiet week breaks in a loud one.

  6. The output is read on a phone

    Not a monthly PDF nobody opens. A dashboard people pull up standing in a field, which sets a different bar for both speed and trust.

Hands at a table holding a phone in daylight, an open notebook and a cold coffee beside it.
The deadline was 12:00 local. Manual entries closed at 12:45, with the team asked to aim for 12:30.

17 days

Built, deployed, broken
by reality, repaired.

Nothing here was designed to completion and then built. Roughly half the final system was built after the gates opened, in production, while it was already reporting.

30 July

Architecture locked

Measurement window defined. Earned-media reach calculated per account rather than as a blended average — in a list holding both a 17-million-follower celebrity and a 3,000-follower micro-creator, an average is a fiction. Paid-advertising connections built immediately and kept on standby until the team defined its KPIs. Certain metrics accepted as permanently manual. First full rehearsal run the same day.

31 July

Live in production

11 collector workflows plus a watchdog, all active. The same day, the earned-media account lists were verified by hand with the team: dead accounts removed, new ones added, and one impostor caught — a lookalike handle impersonating a real public figure, which would otherwise have contributed fabricated reach to the client's own reporting.

31 July, evening

The definition of a day changed

The client moved the measurement window from 16:00→16:00 to 12:00→12:00. That touched every schedule, every window calculation, the watchdog's expectations, the finaliser's day offset, the dashboard subtitle and every page of two team-facing manuals. Rebuilt the same night and verified by a script that asserted no trace of the old window survived anywhere in the system.

1 August

First fully automatic wave

12:05 to 12:45. 9 of 10 collectors succeeded; the 10th failed because the press-monitoring bulletin is not published at weekends. Raw row count exactly equal to the expected sum, zero duplicates. The watchdog emailed “issues, 7 of 10” — and every one of those flags was worth reading.

2 August

An exact follower count, without waiting a month

The platform will not give you your own exact follower number without a formal app review taking one to four weeks. A sandbox route requiring no review turned out to serve real production statistics. 2 readings 3 hours apart differed by exactly 1 follower — which is itself the proof the number was live. Manual entries dropped from 18 to 17.

5–6 August

Target tracking, then hourly sentiment

A live view showing which influencer and creator-army targets were actually producing data and which were silently missing — because a collection failure and genuine inactivity look identical otherwise. Then hourly sentiment, because a single daily mood score during a live event is nearly useless; the interesting signal is the shape of the day.

6 August

Gates open

125k people come through the gates on the first day alone. The system is 3 days old.

7 August, 10:20

The next edition's presale opens, mid-festival

On day 2, with the site already full, fifteen individually tagged launch links went live — profile, story and post per channel, the volunteer creator army, influencers, newsletter, messaging, press, communities and the chatbot. The nine-day-old system began attributing sales by channel while the campaign was still running.

7 August, afternoon

Peak volume breaks the assumptions

Own TikTok output went from ~140 to ~250 videos a day; own Instagram and Facebook to ~47 posts each. Result caps everywhere were silently truncating. Caps raised across six workflows in one afternoon, page cursors added where listings had been clipping, and a new flag added to separate “we hit the ceiling” from “the ceiling actually cost us data.” Cost: about +$76/day, approved as a line item.

9 August

85,457 datapoints in a single day

The busiest day of the run, 2 days into the presale. A nightly evolution wave was added the same night — a job that re-collects days already closed and records every value that changed, so the team can see how much a number moves after it is first reported.

9–15 August

Close, and stop

The festival closed on the 9th. Several jobs self-disable past that date by design: they still fire, do nothing, and write “stopped” to the log. Collection ran on for another six days, and ceased on the 15th. Final total, 466,911 values across 17 days.

Architecture

No warehouse.
No BI licence.
A spreadsheet.

Everything runs on a self-hosted workflow-automation platform. The store of record is a single spreadsheet. That is a deliberate choice, not a compromise: the marketing team already lived in spreadsheets, and putting the data anywhere else would have meant they never looked at it. The cost of that choice is that the spreadsheet becomes a concurrency problem, which it duly did.

The discipline that holds it together: collectors only ever append raw facts in long format — day, source, metric, value, collected at, status. They never write to the presentation grid. The grid is formulas reading the raw tabs. A bad collector run therefore cannot corrupt the dashboard's structure, and any number can be traced back to the run that produced it.

Sources

11 external platforms

social · analytics · ads · email · messaging · press · scraping

Collection

21 workflows · 539 steps

4 tiers: daily 12:05–12:45 · live 6×/day · hourly sentiment · nightly maturation

Supervision

Watchdog, 12:45

asserts every source reported · emails pass/fail

Store of record

Append-only raw tabs

day | source | metric | value | collected_at | status

Presentation

Formula grid + 3 dashboards

HTML served from a webhook, shared-secret key, no build step

Collectors append. Formulas read. Nothing that fetches ever writes to the grid a human looks at.

Tier 1 · once daily

Official

12:05–12:45. The number of record for the closed day. Feeds the leadership grid.

Tier 2 · 6× daily

Live intraday

The day currently in progress. Writes to a separate tab so it can never contaminate the official numbers.

Tier 3 · hourly

Sentiment

Rolling 24-hour comment sentiment, offset from the intraday runs so two jobs never fight over one token.

Tier 4 · nightly

Maturation

Re-collects closed days and logs what changed. Answers “how much does a number move after we first report it?”

Omit, never fake. A failed fetch produces no row, not a zero. A gap is visible and recoverable; a zero is indistinguishable from a real result and quietly poisons every average, trend and total downstream of it, permanently.
Failure isolation at node level. Almost every external call continues on error rather than aborting the run. One platform being down degrades the day to partial instead of losing the other ten platforms.
Two-pass measurement. Provisional at the deadline, final the next day, both kept, each labelled. Speed and accuracy stop being a trade-off once you are willing to report the same number twice.
End-date guards. Anything that costs money has a hard stop date compiled into it. Past that date it still fires, no-ops, and logs that it has stopped — the difference between a project that ends and one that quietly bills forever.
A production cabin behind a stage at four in the morning: a folding table, two open laptops, coiled cables, an empty chair.
Roughly half of the 21 workflows were built after the gates opened, while the system was already reporting.

Volume

1,329 numbers reached
leadership. 466,911 were
collected to defend them.

Roughly 350 datapoints behind every single figure in the deck. That ratio is the whole argument for building rather than buying: anyone can put 1,329 numbers in a spreadsheet — the work is being able to defend each one.

Own-channel post records198,345
Earned-media post records162,165
Hourly sentiment records60,024
Follower snapshots15,824
Daily aggregate metrics — the numbers of record7,980
Boost and promotion tracking3,951
Top-post records3,792
Earned-media account registry — ~470 accounts3,557
Run log — the system's own audit trail2,200
Evolution log — values that changed after first reporting1,941
69,4566 Aug
79,0047 Aug
76,8578 Aug
85,4579 Aug · peak

Most of this exists nowhere else in the world. Not at the platforms, not at the client. A day the collector missed is simply gone.

3,956 follower snapshots exist because most platforms will tell you your follower count now and will never tell you what it was yesterday. 328 evolution records are the system auditing itself — every value that turned out to be different once the data settled.

Failure log

16 things
that went wrong.

The architecture is unremarkable. The failure modes are what nobody writes down. 15 of these 16 were caught before they reached a leadership deck; the one that was not — a dashboard reading a stale range — was caught within days by a person who noticed a number looked old.

Showing all 16

What happened

4 days into the build, a headline reach metric the client had explicitly asked for stopped being available. It had been deprecated in 2 waves, one in late 2025 and one in mid-2026, with no replacement named in the documentation. No error, no warning, no email — the endpoint simply returned nothing for that field.

The fix

Changelog archaeology turned up successor metrics that were not documented as replacements. They were validated against the real account over 2 consecutive days before being trusted. The complication: the replacement buckets data into calendar days ending at 07:00 UTC, and the reporting window was noon-to-noon local. Those never line up.

So the bucket selector refuses to guess. One bucket unambiguously covering the window, use it. Ambiguous, write no row and dump the raw response for a human.

The transferable bit

The danger is not the outage — outages are loud. The danger is the plausible replacement that quietly means something slightly different. Validate any substitute against a period where you already know the answer.

What happened

No free, immediate route to an exact follower count for an account you own. The app shows a rounded figure, and a rounded figure cannot measure daily growth — it does not change for months. The official routes both require app review, one to four weeks. There was no time.

The fix

Sandbox mode requires no review. Whether it serves real production statistics or dummy data was undocumented either way, so it was tested: it serves real data, exactly, in about thirty minutes of setup.

Then a second bug appeared. With exact numbers available, growth is a difference against yesterday's snapshot — and the lookup for “yesterday” was not filtered by account, so it could match any of ~88 tracked handles. It had never yet produced a wrong number, but it would have.

The transferable bit

Half of what an API can do is undocumented, and the documentation will not tell you it is undocumented. Separately: the moment you start computing differences, you have created a class of bug that produces plausible wrong answers rather than errors.

What happened

The brief included messaging broadcast open and click rates. The endpoint that provides read and click data is not supported for business accounts in this region — not rate-limited, not expensive, absent. Broadcast counts and total contact numbers are exposed by no API at all on the tool in use.

The fix

There is none. Four metrics stayed manual permanently. What was done instead: the client was told on day one, with the specific documentation reference, rather than being given optimistic silence until week two. The team manual gained a permanent note explaining why those four values must be typed in every day.

The transferable bit

Scoping what is impossible is real work with real value, and it belongs before anyone builds a slide around a metric. “We cannot get this, here is why, here is what it costs to collect by hand” beats a number that quietly turns out to be wrong.

What happened

On the first live day, seven workflows fired in the same minute. One wrote its audit row, the API returned success, the workflow reported success — and the row was not in the sheet. Its data rows had landed; its log row had not. Concurrent appends are not transactional under contention, and nothing tells you when a write has been silently dropped.

The fix

Stagger everything by two minutes: 12:05, 12:15, 12:17, 12:19, 12:21, 12:23, 12:25, 12:27, 12:29, 12:45, 13:15. Verified live — with the honest caveat written into the notes at the time that this reduces collision probability and does not make the operation atomic. The durable fix is a single dedicated writer that every workflow calls.

The lost run was deliberately not re-executed, because re-running the collector would have duplicated its data rows. A missing audit row is cheaper than duplicate facts.

The transferable bit

Success responses are not proof of persistence. If several jobs write to one shared document, assume contention and design for it — or verify from the destination rather than from the response.

What happened

Press coverage comes from a daily media-monitoring bulletin that arrives by email. On the first Saturday, no bulletin arrived. The collector correctly found nothing, correctly refused to invent rows, and correctly reported a failure — which was, that day, the right behaviour producing an alarming-looking result.

The fix

The collector now selects by exact date first, then falls back to matching a bulletin's stated coverage span, so a Monday multi-day edition fills the weekend days retroactively.

The transferable bit

Every human-operated upstream dependency has a calendar, and it is not your calendar. Find out what it does at weekends and on public holidays before your watchdog emails you a red alert on day two.

What happened

The evening the system went live, the client changed the measurement window from 16:00→16:00 to 12:00→12:00. It touched every schedule, every window calculation, the watchdog's expectation of which day each job should have produced, the finaliser's day offset — because the watchdog now runs before the finaliser, so the newest finalised day is two days back, not one — the dashboard subtitle, and every page of two team documents.

The fix

Rebuilt the same night, then independently verified by a script asserting the new schedules were right, every active flag was right, and zero traces of the old window idiom existed anywhere. Two historical strings in code comments were annotated as historical rather than deleted, so a future reader would not mistake them for live logic.

That verification also surfaced a pre-existing bug: one collector had been commented out of the watchdog's expected-sources list and had therefore never been audited.

The transferable bit

A definitional change is more dangerous than a feature change, because nothing breaks — everything just means something slightly different. And an audit forced by an unrelated change is one of the most reliable ways to find bugs you did not know you had.

What happened

Web analytics backfills for up to 48 hours. A figure pulled at noon for the previous day undercounts, sometimes badly: in one observed case, 14,562 users provisional against 25,969 final. The provisional figure understated the truth by 44%.

The fix

The client was offered the choice explicitly — wait for accurate data, or have everything by 12:30. They chose the deadline. So: a provisional pull at the deadline, clearly marked provisional, and a finalisation job the next day that re-pulls the closed window and overwrites the row in place. The dashboard shows the best available figure at any moment; the history shows which is which.

The transferable bit

Speed and accuracy are not actually a trade-off if you are willing to report the same number twice. What is unacceptable is reporting a provisional number as though it were final.

What happened

On day 2, output volumes exploded — own TikTok from ~140 to ~250 videos a day, own Instagram and Facebook to ~47 posts each, against an off-season baseline of 5–14. Result caps everywhere were returning their configured maximum and silently truncating. Per-post insight fetches covered under half the posts in the window. Page-size limits were being hit with no pagination configured, so only the first page was ever seen.

The fix

Caps raised across six workflows in one afternoon, cursors added where listings had been clipping, timeouts extended — and a diagnostic flag added that separates “we hit the cap” (permanently true at peak, therefore useless as an alarm) from “the cap actually clipped data inside our reporting window” (the real alarm). Every raise was recorded with its off-season value and a walk-back note the same day.

The transferable bit

A saturated limit is not an error; it is a silent truncation, and it looks exactly like a real result. Instrument the difference between hitting a ceiling and the ceiling costing you data — and write down the walk-back the day you raise it, because nobody remembers in September.

What happened

One dashboard read a fixed range of 2,000 rows from an append-only tab. That tab had grown past 14,000 rows, and finished at 30,800. The dashboard was faithfully rendering the oldest 2,000 records — several days stale — and looked entirely normal doing it. No error. The page loads, the charts draw, the numbers are real numbers. They are just the wrong ones.

The fix

Read ranges raised to 60,000 across all dashboards, and the whole class of bug flagged for review. This is the one defect in the log that reached a reader before it was caught — by a person who noticed a number looked old.

The transferable bit

Fixed ranges over growing data are a time bomb with a pleasant face. Any hard-coded limit in a read path needs either a headroom alarm or an explicit “did we hit the edge?” check.

What happened

Sentiment was computed from “the newest 25 comments” per platform. During a festival, with hundreds of posts a day, the newest 25 comments span roughly the last three hours — so the system was reporting the mood of early evening and labelling it the day's sentiment.

Two related bugs sat underneath it: a filter nominally described as a day window was actually applying a seven-day lookback, and the maturation pass was recomputing sentiment from whatever comments it happened to have — in one case four — and overwriting a figure derived from a proper sample.

The fix

Stratified sampling: 25 videos spread evenly across the reporting window rather than the newest 25. The lookback corrected to a true day window. And a minimum-comment threshold, set at 20, below which a matured sentiment value is omitted rather than written.

The transferable bit

Every “newest N” in your stack is a hidden time bias. On a normal day it approximates the day; on your busiest day — the one you actually care about — it approximates the last hour. And put a floor under any percentage: a ratio from four observations should never overwrite one from four hundred.

What happened

The nightly maturation pass — which exists to update a day's figures once the data has settled — ran while a result cap was clipping its input. It computed zero where the true figures were 175,301 views and 10,165 engagements, and wrote those zeros over the correct values.

The fix

The affected rows were identified and restored from history. A guard was added: the maturation pass may update a value upward, but may never write a zero over a non-zero. If it computes zero, it omits.

The transferable bit

The job that fixes your data is the most dangerous job you have, because it has write access to numbers that were previously correct. Give every correcting process an explicit refusal condition, and make it monotonic wherever the underlying quantity can only grow.

What happened

Post data was joined to the tracked account list by the post's owner username. On investigation, 29% of relevant posts were collaborative posts owned by a co-author — so they appeared under a different account and dropped out of the tracked account's totals entirely.

The fix

Join by the profile that was requested, not by the owner reported on the returned post.

The transferable bit

Join on what you asked for, not on what came back. Every social platform has some feature — collabs, reposts, cross-posts, shared reels — that makes “who owns this post” a different question from “whose feed did we find it in.”

What happened

Manual verification of the earned-media lists surfaced a catalogue of individually small problems: accounts with no online presence listed as if they had one; an impostor account impersonating a real public figure, which would have contributed fabricated reach to the client's own reporting; age-gated accounts that cannot be scraped at all; an account dormant on one platform and active on the other; handles changed since the list was compiled; and the same person in two groups, which is fine but means group totals are not additive.

The fix

A verification pass with the team, a per-account explicit track / don't-track flag so an exclusion is recorded rather than implied by absence, a confidence column, and a live view showing which targets were producing no data so gaps stayed visible.

The transferable bit

Influencer and partner reporting is a data-quality problem before it is a data-collection problem. Budget real time for list verification, and never let a list be maintained purely by absence.

What happened

The programmatic interface for editing workflows had a broken partial-update operation, could not activate a workflow at all, and choked on payloads above roughly 40KB — while several of these workflows are 40–170KB of JSON. One workflow's update schema rejected a field the server itself had set, so it had to be stripped from every update.

The fix

Custom deployment tooling: assertion-guarded find-and-replace against local dumps of each workflow, then full replacement via the API, with verification scripts that read the live state back and assert the change landed.

The transferable bit

Meta-work is work. When a platform's own editing tools cannot handle the size of the thing you have built, a small safe deployment path is not a distraction — it is the only thing that lets you keep changing a live system without fear.

What happened

On the first live day the watchdog emailed “ISSUES: 7 of 10.” Investigation showed one false positive from the lost-write race, one first-day artefact for a job whose first execution had not happened yet, one genuine weekend bulletin gap, and “17 of 17 manual entries missing” — which was correct, because the team had not yet started their daily routine.

The actual data outcome that day: nine of ten collectors fine, raw row count exactly equal to the expected sum, zero duplicates.

The fix

None needed. Each alert was understood before anything was tuned.

The transferable bit

A watchdog that reports problems on day one is working. Tune it after you have understood each alert, not before — and resist the urge to soften alerts to make the email look better. The uncomfortable email is the product.

What happened

When the presale opened mid-festival, the obvious query — sales for the new edition — returned nothing. Not an error. Zero rows. In the ticketing system the new passes carry the outgoing edition's identifier, which is reasonable enough from the shop's point of view and fatal from a reporting one: the single most obvious field to filter on is the wrong field.

The second-order damage runs the other way too: a straightforward query for the outgoing edition silently swallows the entire presale. Two reports, wrong in opposite directions, from one modelling decision — and both returning confident, plausible numbers.

The fix

Filter on the ticket name, never on the edition field, documented at the top of every query template as the first thing anyone touching this data must read. A related trap in the same workstream: the shop sells both editions at once, so session-scoped channel reports attribute one edition's revenue to the other's campaigns. Item-scoped metrics filtered on product name are the only correct approach.

The transferable bit

The most dangerous data is not missing or malformed — it is correctly-formatted data sitting in a category you did not expect. It passes every validation and hands you a number you have no reason to doubt. Run the query that should return zero and check that it does.

Platform by platform

What each one gives you,
and what it takes back.

Established by building against these APIs in July and August 2026, not from documentation. Platforms change — treat the specifics as dated and the patterns as durable.

Facebook and Instagram

Works wellPost listings, per-post engagement, follower counts, story listings, comment retrieval, page and account resolution. Long-lived system-user tokens are strongly preferable to anything a human has to re-authorise mid-event.

  • Deprecated reach. The classic reach metric family was removed in two waves. Replacements exist under different names and are not documented as replacements — and one remaining impressions metric is itself on the published deprecation list.
  • Bucket alignment. The replacement page-level metric returns calendar-day buckets ending at 07:00 UTC. If your reporting day is anything else, the mapping is ambiguous and must be handled explicitly.
  • Silent page-size saturation. Listing endpoints accept a limit of 100 and will happily return exactly 100 with no indication there were 250. At peak this was the difference between seeing half the posts and all of them.
  • Collaborative posts. 29% of relevant posts here were owned by a co-author. Joining by the returned owner loses them.

Silent deprecation. Silent truncation. Silent immaturity. None of the three throws an exception, and all three produce numbers a dashboard renders beautifully.

The outcome

The launch was measured
while it was happening.

The presale for the next edition opened on day 2, at 10:20, with 125k people on site. It was instrumented end to end before it opened: fifteen individually tagged links, per-channel sales attribution joining the ticketing database against web analytics with an explicit reconciliation line, and paid campaigns sitting in the same revenue-share table as the organic ones.

A channel that was underperforming — or that had silently never been published at all — could be named and fixed inside the selling window, instead of appearing in a post-mortem.

What is evidenced

  • The launch was fully instrumented before it opened.
  • 40,000 passes sold.
  • Channel-level sales attribution was available during the presale window, not weeks after it.
  • Because every link was tagged individually, a link with zero traffic was identifiable as never published rather than published and ineffective — completely different problems with completely different fixes. In practice this was the single most actionable output of the launch reporting.
  • The infrastructure carrying it was the same system built four days before the gates opened, running through the busiest days it would ever see.

What is not, and will not be claimed

  • That the monitoring system caused any specific number of ticket sales. There is no counterfactual — nobody ran this launch without instrumentation to compare against.

Published with UNTOLD's agreement. 40,000 passes sold. Attendance is not a measurement taken by this system. Everything describing the system itself — workflow counts, datapoints collected, metrics, costs — is as recorded in the production systems on 17 August 2026, eight days after the festival closed.

What it cost to run

Under $150 a day,
at peak.

Item Normal Festival peak
Third-party scraping — earned media and own content ~$17–18/day ~$130/day
LLM calls — sentiment classification, daily narrative negligible negligible
Analytics, ads, email and social APIs $0 $0
Workflow automation platform self-hosted, already in place
Storage, warehouse, BI licensing $0 $0
Total tooling ~$20/day under $150/day

Full-festival scraping for the earned-media layer landed around $200–250, plus the peak-period raises. Against 466,911 datapoints, that is roughly $0.004 per datapoint. A late addition — a nightly re-collection wave that re-fires the expensive collector for each closed day — was costed at ~$31/night and ~$124 over four nights, and approved as a line item before it was switched on.

What is not in these numbers: labour. This is tooling cost only. Presenting it as a project price would be dishonest, and it is not a quote for similar work.
Every cap raise recorded its walk-back the same day. “Raised from 12 to 30 for peak; return to 8–12 after 15 August.” Without this, festival-tuned settings run at festival prices against off-season volumes forever, and nobody notices until an invoice looks odd in November.
Costed before switched on. The nightly re-collection wave was priced per night and approved explicitly, rather than discovered on a bill.

The security posture, honestly

Building at this speed, under this deadline, produced real security debt: credentials with unrestricted outbound-domain permissions rather than allow-lists; API tokens that passed through chat during debugging; at least one third-party secret stored in plaintext inside a workflow definition rather than in the credential store; and dashboards protected by a shared secret in a URL query parameter — adequate against casual discovery, not against a leaked link.

All of it was written down as it was incurred, in a single loose-ends list with a stated remediation: rotate every credential, move plaintext secrets into the credential store, rotate the dashboard keys, and deactivate the collection layer after the final report. Speed has a cost and it is usually paid in security posture. The professional move is not to pretend otherwise — it is to keep an explicit debt register from day one, so the remediation is a task list rather than an archaeology project.

A person seen from behind at a desk before dawn, facing two monitors, the window beyond showing the first grey of morning.
Manual entries went from 18 a day to 16. Every reduction was proven live in production before the team's documentation was changed, so no value was ever entered by both a human and a machine.

What transfers

Eleven things worth
taking somewhere else.

  1. A gap beats a guess

    Build every collector so failure produces no value — never a zero, never an estimate. A gap is visible and recoverable. A zero is indistinguishable from a real result and poisons every average downstream of it, permanently, in a way nobody will ever find.

  2. Assume your metrics will be deprecated

    Know which of your reported metrics are on a deprecation list right now. When you accept a substitute, validate it against a period where you already know the answer.

  3. Measure twice, and say which one you are showing

    Report provisionally at the deadline, finalise the next day, keep both, label them. Anyone who has had to explain to a board why Tuesday's number changed will understand why the extra job is worth it.

  4. Name the impossible on day one

    The instinct to keep looking rather than deliver bad news is the expensive one. Saying “we cannot get this, here is the documentation” stops a strategy being built on a metric that will never arrive.

  5. Something has to watch the watcher

    A dashboard with a missing number still looks like a dashboard. Run an independent daily audit that asserts every expected source reported — and resist softening it when the email is uncomfortable.

  6. Every “newest N” is a hidden time bias

    Sampling the newest 25 comments approximates the day on a quiet day and the last hour on your busiest one. Stratify across the window, then put a floor under every percentage.

  7. A saturated limit is a silent truncation

    Hitting a cap does not throw an error; it returns a valid, plausible, incomplete answer. Treat every hard-coded read range over growing data as a time bomb, because it is.

  8. Success responses are not proof of persistence

    Concurrent writes to a shared document can be accepted, confirmed and dropped. Serialise them through a single writer, or verify from the destination rather than the response.

  9. The job that corrects your data is your most dangerous job

    It is the only thing with write access to numbers that were previously right. Give every correcting process an explicit refusal condition.

  10. Write down the walk-back the day you raise the limit

    Every cap raised for a peak needs its off-season value recorded immediately, with a date. The same goes for end-dates on anything that costs money: compile the stop date in rather than trusting anyone to remember.

  11. Verify the field before you trust the filter

    Correctly-formatted data in the wrong category throws no errors and returns instantly. Before trusting any filter on a categorical field, run the query that should return nothing and confirm that it does. Two extra queries. Ninety seconds.

AI did not remove the hard part. It removed the slow part.

The hard parts here were judgement: discovering a metric had been deprecated, reasoning about whether a sandbox API would serve production data, deciding a missing number beats a wrong one, noticing that “newest 25 comments” is a time bias, and choosing to tell the client on day one that four metrics were impossible.

What collapsed was everything around them — reading changelogs across a dozen platforms in an afternoon, writing several hundred nodes of integration logic, regenerating the team documentation six times as the scope moved, building custom deployment tooling because the platform's own could not handle the payload size.

The practical consequence for anyone commissioning this kind of work: the cost structure has inverted. Implementation used to be expensive and specification cheap. Now specification — knowing what to ask for, what is impossible, and what a number actually means — is the expensive part. Which means the right question to ask a vendor is no longer “how long will it take to build?” It is “what will you refuse to report, and why?”

Let's talk

Your reporting problem
is not a budget problem.

The tools cost less than one agency invoice. What is scarce is knowing which platform quietly deprecated your metric last quarter, which one lies by omission, and where to put the guardrail that stops a bad number reaching your board.