Skip to content
Tweet Cruncher

Methods · how the numbers were made

Methods, decisions and limits

Where every number on this site comes from, how the 2023 program and its 2026 port were checked, how the browser benchmark and the AI feature are evaluated, and what all of it cannot tell you.

Data

Data provenance

No course data is hosted or re-analysed here. Each artefact below says where it came from and whether it can be regenerated.

Swipe the table sideways for the sources.

ArtefactSourceRe-runnable?
Task 1 to 3 results, dataset summaryTranscribed from the 2023 submission (18.7 GB, 9,092,274 tweets). Task 1 author ids are as printed, with trailing digits lost to a spreadsheet.No: the course data lived only on Spartan.
Spartan benchmark jobsFinal jobs 46094405 to 46094407 from the submission; earlier jobs 45983020 to 45983022 from the Slurm logs on the Spartan-running-test branch. Wall clock in whole seconds, one run per layout.No.
Place gazetteer (117 keys)The original process_salV1 run on sal.json by scripts/build_gazetteer.py, restricted to keys the demo can hit.Yes, with the course's sal.json.
Synthetic tweet filesSeeded generator with bigTwitter.json's line layout (DR-003).Yes: same seed, same bytes.
Parity fixturesThe original Python run outside Spartan on synthetic files by scripts/run_original.py.Yes, with uv and sal.json.
Browser benchmark samplesYour own runs in the MPI lab; exported as CSV or JSON.Yes, on your machine.

Method

The program and its port

The 2023 program splits the file into equal byte ranges, one per MPI rank (DR-001); each rank scans its range line by line with three regular expressions, matches place names to Greater Capital Cities through n-grams of the normalised name, and reduces its tweets to three count tables, which are sent to ranks 0, 1 and 2 for the final answers (DR-002). The step-by-step version is on How it works.

The 2026 revival ports the algorithm to TypeScript, quirks included, and runs it with one Web Worker per rank. The original Python is run outside Spartan on the same synthetic inputs to produce reference outputs, and the test suite requires the port to match them: per-tweet records and per-rank counts for 1, 3, 4 and 7 ranks, all four result files, the boundary double count at 27 and 39 ranks, and the UTF-8 decode crash at 17 and 29 ranks.

Evaluation

Evaluation design

Correctness

Correctness is tested, not sampled: the parity tests above are exact comparisons, and every run in the lab is checked against the first run on the same file that counted no tweet twice.

Performance

Spartan, 2023. Three layouts, one run each, timed in whole seconds. The 8-core speedup was 6.54×, and Amdahl's serial fraction fitted to the final runs is 3.18%. Because both multi-core layouts used 8 cores, the fit is the Karp–Flatt value at n = 8, a point estimate: there is no repeat from which to estimate run-to-run spread, so no interval can be given. Rounding to whole seconds alone moves it between 3.08% and 3.28%.

Browser benchmark, 2026 (DR-004). 1 warm-up round discarded, then 10 timed rounds by default; every worker count once per round, in an order shuffled with seed 2023. Reported per worker count: the median wall time with its exact order-statistic interval, which needs no assumption about the shape of the noise and whose true coverage is shown (97.9% at the default 10 rounds); then speedup (ratio of medians), efficiency and the Karp–Flatt fraction, each with a percentile bootstrap interval (2,000 resamples of whole rounds, seed 90024). Amdahl's f is a least-squares fit to the median speedups, with its interval from the same resamples. Bootstrap intervals are labelled nominal 95%: a seeded simulation put their coverage at 94% to 95% with 5 or 7 rounds and 96% to 98% with 10 or 15, and no interval is shown with fewer than 5 complete rounds. Gustafson's law (n − f(n − 1)) is drawn with the same f for contrast: it assumes the input grows with n, whereas the lab keeps it fixed.

The statistics are unit-tested against values computed independently with numpy, scipy and statsmodels (scripts/stats_reference.py, which replays the seeded bootstrap draw for draw) and with base R (scripts/stats_reference.R). The coverage simulation is web/src/lib/lab/coverage.ts, run with pnpm coverage-sim.

The optional AI feature

Ask the results is evaluated on 24 fixed questions (14 answerable, 10 not) with an answer key computed in code from the same tables; numbers the question already contains do not count as evidence. Every item a run reached is scored, so a call that fails on an item counts as a fail instead of dropping out of the denominator. Rates, including the call-error rate, are reported with Wilson 95% intervals; two runs can be compared item by item with a paired bootstrap interval and an exact McNemar test, with a warning when their request settings differ (DR-005). No scores are published here: running it costs money on a visitor's key, and a score the author ran and reported would be a self-scored metric.

Assumptions

What the analysis assumes

  • The transcribed results match what the 2023 program wrote; they cannot be re-derived.
  • Amdahl's law describes this program's strong scaling: a constant serial part plus work that divides evenly over n workers.
  • In the browser, a Web Worker per rank on a machine with at least that many logical cores behaves like an MPI rank on its own core, and the page relaying tables stands in for MPI send and receive.
  • Benchmark rounds are exchangeable: conditions within a session do not change in a way the shuffled order cannot spread out.
  • For the AI feature, the tables are the whole truth: a question they cannot answer should be declined, even if the model knows the answer elsewhere.

Limitations

What it cannot tell you

  • Spartan's serial fraction is a single point at n = 8. It cannot distinguish Amdahl's law from any other curve through that point, and its ceiling of about 31× is an extrapolation.
  • The browser is not Spartan: memory instead of a shared file system, no network, a browser scheduler, logical cores that include hyper-threads or efficiency cores, and a synthetic file with shorter records and fewer tweets per author (DR-003). The lab also runs the three reductions in parallel, which the 2023 code did not (DR-002).
  • The intervals describe spread on one machine in one session, assuming runs are independent. Between sessions the same machine varied more than any one interval (DR-004).
  • Two 2023 bugs are kept on purpose for fidelity: a chunk boundary inside an "_id" line's indentation double-counts a tweet (about 1 in 515 per boundary on bigTwitter.json), and a boundary inside a multi-byte character crashes a rank. Whether the double count affected the 2023 runs is unknown.
  • The AI evaluation is small (24 items), written by the same person as the prompt, and graded by literal number matching; it is a check on grounding behaviour, not a model ranking.

Next time

What I'd change

  • On Spartan: repeat each layout at least five times, add 2, 4 and 16 cores, and log the time of every phase on every rank, so the serial fraction is measured with an interval instead of inferred from one point.
  • Snap chunk boundaries to record starts, which removes both boundary bugs (DR-001), and post the three gathers as non-blocking sends so the reductions really overlap.
  • Calibrate the synthetic data to the published ratios (about 76 tweets per author and 2 kB per tweet) so the lab's balance of scanning and reducing is closer to Spartan's.
  • Fit and compare a model with a communication term that grows with n, and show residuals.
  • Have someone else write a held-out question set for the AI evaluation.

Decision records

Why it was built this way

One record per decision, in a fixed format: context, decision, options considered, why, what happened (weak numbers included) and what I'd change. Records are never edited after the fact; a later record supersedes an earlier one.

  1. DR-001Split the file by byte ranges, not by linesGive each MPI rank an equal byte range of bigTwitter.json (ceil(file_size / ranks) bytes), let it seek straight to its offset and scan lines from there, and never count or index lines before the parallel work starts.Read the record
  2. DR-002Pre-aggregate on every rank, then send the tables to three task ranksEach rank reduces its own tweets to three small count tables before any communication (per author, per capital city, per author and city), then sends each table point to point to the rank that owns that task: rank 0 for Task 1, rank 1 for Task 2, rank 2 for Task 3. With one rank, rank 0 does all three.Read the record
  3. DR-003A seeded synthetic file with bigTwitter.json's exact line layoutThe browser demo and the parity tests run on a generated file that copies the line layout of bigTwitter.json exactly, filled with made-up ids, text and places from a seeded generator, rather than on any real tweets.Read the record
  4. DR-004Measure browser scaling with repeated, shuffled rounds and intervals of checked coverageThe MPI lab's benchmark runs every worker count R times (10 by default; 5, 7, 10 or 15 offered) after a discarded warm-up round, in a seeded random order within each round. It reports each median wall time with the exact order-statistic interval, and speedup, efficiency, the Karp–Flatt fraction and Amdahl's serial fraction with percentile bootstrap intervals that resample whole rounds, labelled "nominal 95%" because their real coverage was checked by simulation. With fewer than 5 complete rounds it shows medians and ranges only. Amdahl's f is fitted by least squares to the median speedups, and Gustafson's law is drawn for contrast.Read the record
  5. DR-005Answer questions from the published tables only, with citations and an audit trail"Ask the results" sends the visitor's question, with the transcribed result and benchmark tables as the only context, straight from their browser to the provider of their choice using their own key. The model must cite the row ids it used, write out any arithmetic, and say "not answerable" when the tables do not cover the question. Every answer is checked automatically, labelled as AI-generated, left for a person to accept, edit or reject, and logged in the browser with the SHA-256 of the context. A fixed 24-question evaluation with an answer key computed from the data measures how well a chosen model keeps to these rules.Read the record

Model card

Intended use, data, evaluation with its limits, failure modes and ethical considerations for the Amdahl scaling model and the grounded question answering.

Read the model card

Transparency

AI use statement

This site has one optional AI feature, Ask the results, and an evaluation harness for it. Everything else on the site, including the results, the scaling analysis, the MPI lab and its benchmark, works without AI and without an API key. The design is informed by the Australian Government's policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It is not a claim of compliance with any of them.

What the AI does

  • Answers questions about the published 2023 results and the Spartan benchmark jobs, using only those tables as context.
  • Cites the table rows it used, writes out any arithmetic, and says when a question cannot be answered from the tables.
  • In the evaluation harness, answers a fixed set of 24 questions so its answers can be graded against an answer key computed from the data.

What it never does

  • It does not produce or change any result shown elsewhere on the site. The original results are transcribed from the 2023 submission and the benchmark analysis is computed in code.
  • It does not see the raw tweets. They are course data and are not on this site.
  • It is not asked to identify, profile or speculate about the people behind the author ids in the results, and it is instructed to refuse.
  • It does not act on its own: nothing it writes is saved anywhere except your browser's audit log, and only after you ask.

What is sent, and to whom

  • To the provider you choose (Anthropic or OpenAI): a fixed system prompt containing the result tables (shown in full on the Ask page, with their SHA-256 hash) and your question. Calls go directly from your browser to the provider and are billed to your own key under the provider's terms.
  • Your API key is stored in your browser's session storage (or local storage if you tick "remember on this device") and sent only to the provider. This site is static: there is no server of ours that could receive the key, the question or the answer. The key is never written to the audit log or included in exports.
  • Nothing is sent to the site's author.

Human in the loop

Every answer is labelled AI-generated, shown with the rows it cites (taken from the site's own data, not from the model's text) and with the result of automatic grounding checks, which include re-doing the arithmetic the model shows. You decide whether to accept it, edit it or reject it, and every decision is recorded in order, so an edit is kept even if you later accept or reject the answer. Evaluation answers are graded automatically against the answer key and are marked as such.

Audit trail

Every call, successful or not, is appended to an audit log in your browser (IndexedDB): time, feature, provider, model requested and model that answered (recorded separately), the request settings (output-token limit, effort, server-side fallback, a hash of the full system prompt), the prompts sent (never the key), the context hash, the output or the error and any raw reply, latency, token usage when the provider reports it, and your decisions. If the browser cannot write the log (for example, storage is full), the answer is still shown, marked as not logged. Text cells in the CSV export that a spreadsheet would treat as a formula are prefixed with an apostrophe. You can read, export (JSON or CSV) and clear it at /ai-log.

Models

Anthropic is the default provider, with Claude Haiku 4.5 as the default model (lowest cost) and Claude Sonnet 5.5 as an option. Sonnet 5.5 requests opt in to Anthropic's server-side fallback, so if it declines on safety grounds another Claude model may answer; the audit log records which one did. With OpenAI you choose the model id.

Limitations

Models can still misread a row, make arithmetic mistakes, or answer when they should decline; the checks catch some of this, not all of it. The evaluation set is small (24 questions) and was written by the site's author, so it is a check, not a benchmark. Results differ between models and between runs.

How this site was built

The 2026 revival's code and documents were produced with the help of an AI coding assistant; the commits that it co-authored carry a Co-Authored-By trailer.