How it works
One file, many ranks, three small tables
bigTwitter.json was too big to load into memory and too slow to parse as JSON on one core. The program never parses it: every MPI rank reads only its own slice of bytes, pulls out three fields with regular expressions, and ships tiny count tables back for the final answer.
- 01
Split
Cut the file into equal byte ranges, one per rank.
- 02
Scan
Each rank seeks to its offset and line-scans with three regexes.
- 03
Match
Normalise each place name and look its n-grams up in sal.json.
- 04
Count
Group-by on the rank: per author, per city, per author-city pair.
- 05
Gather
Send the small partial tables to ranks 0, 1 and 2.
- 06
Reduce
Sum, rank and write task1.csv, task2.csv and task3.csv.
01 · split_file_into_chunks
Byte ranges, not lines
Counting lines would mean reading the whole file first, so the program divides its byte size instead. A boundary will usually land in the middle of a tweet, and two rules make that safe:
- A rank keeps reading past its
chunk_enduntil its last tweet has a place, so it always finishes what it started. - A rank that starts mid-tweet never saw that tweet's
_id, so the author and place lines that follow are ignored: its counters are already level.
So every tweet is counted by exactly one rank, with one exception we only found while porting. If a cut lands inside the four spaces of indentation before "_id", the next rank's partial first line still matches the regex, and both ranks count that tweet. At about 2 kB per tweet that is a 1 in 500 chance per boundary on bigTwitter.json. The original Python does it too, so the port keeps the bug and a parity test pins it; the lab flags it when one of your runs hits it.
| Rank | chunk_start | chunk_end | Size | Task host |
|---|---|---|---|---|
| 0 | 0 | 2,341,913,383 | 2.34 GB | task 1 |
| 1 | 2,341,913,383 | 4,683,826,766 | 2.34 GB | task 2 |
| 2 | 4,683,826,766 | 7,025,740,149 | 2.34 GB | task 3 |
| 3 | 7,025,740,149 | 9,367,653,532 | 2.34 GB | – |
| 4 | 9,367,653,532 | 11,709,566,915 | 2.34 GB | – |
| 5 | 11,709,566,915 | 14,051,480,298 | 2.34 GB | – |
| 6 | 14,051,480,298 | 16,393,393,681 | 2.34 GB | – |
| 7 | 16,393,393,681 | 18,735,307,060 | 2.34 GB | – |
chunk_size = ceil(file_size / size); the last range absorbs the remainder. Offsets are for the real 18,735,307,060-byte bigTwitter.json.
02 · twitter_processorV1
A line scanner with magic numbers
The file is pretty-printed JSON, one field per line, with a fixed layout. After finding a tweet's _id the scanner skips 2 lines; after author_id it skips 18; after full_name it skips 20. Those counts (commented # MAGICS NUMBERS in the original) jump over text and metadata without even running a regex on them. Below is the real trace of the ported scanner over the first two tweets of a synthetic file.
27 lines regex-tested, 80 skipped unread
- 1read:[
- 2read: {
- 3read: "_id": "1464758956051061163",_id captured
1464758956051061163; skip the next 2 lines - 4skipped: "_rev": "2-5c5d032c9bb4fa870793cd785e77ad94",
- 5skipped: "data": {
- 6read: "author_id": "589236440400667814",author_id captured
589236440400667814; skip the next 18 lines - 7skipped: "conversation_id": "1464758956051061163",
- 8skipped: "created_at": "2021-11-28T00:52:03.684Z",
- 9skipped: "geo": {
- 10skipped: "place_id": "af7c5ad704ea233b"
- 11skipped: },
- 12skipped: "lang": "en",
- 13skipped: "public_metrics": {
- 14skipped: "retweet_count": 0,
- 15skipped: "reply_count": 0,
- 16skipped: "like_count": 0,
- 17skipped: "quote_count": 0
- 18skipped: },
- 19skipped: "text": "Back home in Newcastle for the first time in ages.",
- 20skipped: "sentiment": 0.014230137690902
- 21skipped: },
- 22skipped: "includes": {
- 23skipped: "places": [
- 24skipped: {
- 25read: "full_name": "Newcastle, New South Wales",full_name → normalised
“newcastle nsw”, first sal_dict hit“newcastle”→ 1rnsw (Rest of NSW); skip the next 20 lines - 26skipped: "geo": {
- 27skipped: "type": "Feature",
- 28skipped: "bbox": [
- 29skipped: 151.653661,
- 30skipped: -33.049339,
- 31skipped: 151.898339,
- 32skipped: -32.804661
- 33skipped: ],
- 34skipped: "properties": {}
- 35skipped: },
- 36skipped: "id": "af7c5ad704ea233b"
- 37skipped: }
- 38skipped: ]
- 39skipped: },
- 40skipped: "matching_rules": [
- 41skipped: {
- 42skipped: "id": 1412189062442586000,
- 43skipped: "tag": "synthetic: geotagged tweets from Australia"
- 44skipped: }
- 45skipped: ]
- 46read: },
- 47read: {
- 48read: "_id": "1516320407265348431",_id captured
1516320407265348431; skip the next 2 lines - 49skipped: "_rev": "1-168b703c45618c8597c915b99d5042f9",
- 50skipped: "data": {
- 51read: "author_id": "589236440400667814",author_id captured
589236440400667814; skip the next 18 lines - 52skipped: "conversation_id": "1516320407265348431",
- 53skipped: "created_at": "2022-04-19T07:38:51.619Z",
- 54skipped: "entities": {
- 55skipped: "hashtags": [
- 56skipped: {
- 57skipped: "start": 49,
- 58skipped: "end": 54,
- 59skipped: "tag": "data"
- 60skipped: }
- 61skipped: ],
- 62skipped: "urls": [
- 63skipped: {
- 64skipped: "start": 55,
- 65skipped: "end": 78,
- 66skipped: "url": "https://t.co/054ea03b7e",
- 67skipped: "expanded_url": "https://twitter.com/i/web/status/1516320407265348431",
- 68skipped: "display_url": "twitter.com/i/web/status/151632..."
- 69skipped: }
- 70read: ]
- 71read: },
- 72read: "geo": {
- 73read: "place_id": "af7c5ad704ea233b"
- 74read: },
- 75read: "lang": "en",
- 76read: "public_metrics": {
- 77read: "retweet_count": 0,
- 78read: "reply_count": 0,
- 79read: "like_count": 32,
- 80read: "quote_count": 1
- 81read: },
- 82read: "text": "Footy tonight in Newcastle - who else is around? #data https://t.co/054ea03b7e"
- 83read: },
- 84read: "includes": {
- 85read: "places": [
- 86read: {
- 87read: "full_name": "Newcastle, New South Wales",full_name → normalised
“newcastle nsw”, first sal_dict hit“newcastle”→ 1rnsw (Rest of NSW); skip the next 20 lines - 88skipped: "geo": {
- 89skipped: "type": "Feature",
- 90skipped: "bbox": [
- 91skipped: 151.662001,
- 92skipped: -33.040999,
- 93skipped: 151.889999,
- 94skipped: -32.813001
- 95skipped: ],
- 96skipped: "properties": {}
- 97skipped: },
- 98skipped: "id": "af7c5ad704ea233b"
- 99skipped: }
- 100skipped: ]
- 101skipped: },
- 102skipped: "matching_rules": [
- 103skipped: {
- 104skipped: "id": 1412189062442586000,
- 105skipped: "tag": "synthetic: geotagged tweets from Australia"
- 106skipped: }
- 107skipped: ]
03 · normalise_location + sal_dict
From a place name to a capital city
sal.json maps 15,340 suburb and locality names to Greater Capital City codes such as 2gmel, or rural codes such as 2rvic. The program cleans the keys (brackets, “ - ” and full stops removed), adds every word pair of names longer than two words, and builds a 16,616-key dictionary.
Each tweet's place is lower-cased and normalised, then every combination of its words is tried, shortest first. The first key that exists wins. It is fast and usually right, and the port keeps its quirks: try Macquarie Park, Sydney.
- 1
Lower-case
“box hill, melbourne” - 2
normalise_location: drop punctuation, abbreviate state names, squeeze spaces
“box hill melbourne” - 3
Try every word combination in itertools order (shortest first) against sal_dict; the first hit wins
box (no match)hill (no match)melbourne (match)box hillbox melbournehill melbournebox hill melbourne3 words → 7 combinations; the first hit was number 3.
- 42gmelGreater Melbournecounts towards Tasks 2 and 3
Verified: the original Python resolved this exact place to 2gmel against the full sal.json.
04 · gather_task_tdf + reductions
Gather onto three task ranks
Each rank turns its tweets into three partial tables. Rather than sending everything to rank 0, the program spreads the reductions: rank 0 hosts Task 1, rank 1 Task 2 and rank 2 Task 3. Each host receives its own table first, then one from every other rank in order, and writes one CSV.
Task 1 on rank 0
author_id → tweet count
Sum counts per author, rank with ties sharing the best place (method="min"), keep rank ≤ 10. Ties can therefore produce more than ten rows.
Task 2 on rank 1
gcc → tweet count
Drop unmatched tweets and rural codes (\dr[a-z]{3}), sum per capital city, sort by code. “Other Territories” (9oter) is not rural, so it stays.
Task 3 on rank 2
(author_id, gcc) → tweet count
Count distinct cities per author, order by cities then tweets, take the first ten, and spell out the per-city counts, for example #1879gmel.
Fidelity
Ported, not reinvented
The browser runs a line-for-line TypeScript port of the 2023 code, magic numbers and quirks included. The original Python in coursework/ (analysis logic as submitted) was run outside Spartan (with mpi4py stubbed and the ranks re-enacted in order) to produce reference outputs, and the test suite checks the port against them:
In the test suite (runs in CI)
Fixtures produced by the original Python
- per-tweet records (id, author, normalised place, gcc) identical
- per-rank tweet counts identical for 1, 3, 4 and 7 ranks
- task1.csv, task2.csv, task3.csv and task3_1.csv identical
- process_salV1 on a sal.json sample: same keys, codes and insertion order
- every demo place resolves exactly as it does against the full sal.json
Checked locally with the course files
Not committed: course data stays off GitHub
- the full processed sal_dict: all 16,616 keys and codes identical
- tinyTwitter.json: records, per-rank counts and all four CSVs identical for 1, 3, 4 and 8 ranks