Skip to content
Tweet Cruncher

MPI lab · runs entirely in your browser

Run the 2023 pipeline on your own cores

Each Web Worker plays one MPI rank. The file is cut into byte ranges with the original split_file_into_chunks, every rank runs the TypeScript port of twitter_processorV1 over its share, and ranks 0, 1 and 2 reduce Tasks 1, 2 and 3, exactly as main.py did on Spartan. New here? See how it works.

1

Input file

A seeded, made-up file with the exact line layout of bigTwitter.json, so the original scanner's skip counts still line up. Same seed, same bytes.

Tweets

About 122 MB, built in a background worker and kept in memory as a Blob.

sal_dict loading gazetteer.json

2

MPI ranks

detecting cores
2 ranks: unsupported by the original

The file is cut into 8 byte ranges. Task 1 is reduced on rank 0, Task 2 on rank 1, Task 3 on rank 2.

Benchmark settings

Benchmark

Measures every worker count 10 times after 1 warm-up round, each round in a shuffled order, and reports medians with exact order-statistic intervals and bootstrap intervals for speedup and f. At least 5 rounds; 10 keeps the bootstrap at its nominal 95% in simulation.

Repeats
Warm-up
Counts

n = 1, 3, 4, 6, 8 · (1 + 10) × 5 = 55 runs

3

Rank monitor

Generate or pick an input file, then run it. Each rank will appear here with its byte range, live progress and, afterwards, a timeline of scan and reduce phases.

4

Scaling on this machine

Run the benchmark to measure speedup on this machine with uncertainty: every worker count is run several times (10 by default) after a warm-up round, in a shuffled order each round, and the results are summarised as medians with exact order-statistic intervals, speedups and an Amdahl fit with bootstrap intervals, and Gustafson's law for contrast.

Why Spartan's f has no interval. The 2023 benchmark is 3 configurations × 1 run, and both multi-core jobs used 8 cores, so its serial fraction (3.2%) is the Karp–Flatt value at n = 8: a point estimate with no repeat to show run-to-run spread. Whole-second Slurm timing alone moves it between 3.08% and 3.28%. The benchmark here repeats every configuration so it can report an interval. Method

Answers from the task ranks

The three answers appear here once a run finishes.

Runs and output checks

Finished runs on the current file and place dictionary are listed here, each checked against the first run that counted no tweet twice.

The port is checked against the original Python in the test suite: on the same input, per-tweet records, per-rank counts and all four result files match for 1, 3, 4 and 7 ranks. Timings here are wall-clock in your browser and include the page relaying partial tables between workers, which stands in for MPI's send and receive.