Skip to content
Tweet Cruncher

How it works

One file, many ranks, three small tables

bigTwitter.json was too big to load into memory and too slow to parse as JSON on one core. The program never parses it: every MPI rank reads only its own slice of bytes, pulls out three fields with regular expressions, and ships tiny count tables back for the final answer.

  1. 01

    Split

    Cut the file into equal byte ranges, one per rank.

  2. 02

    Scan

    Each rank seeks to its offset and line-scans with three regexes.

  3. 03

    Match

    Normalise each place name and look its n-grams up in sal.json.

  4. 04

    Count

    Group-by on the rank: per author, per city, per author-city pair.

  5. 05

    Gather

    Send the small partial tables to ranks 0, 1 and 2.

  6. 06

    Reduce

    Sum, rank and write task1.csv, task2.csv and task3.csv.

01 · split_file_into_chunks

Byte ranges, not lines

Counting lines would mean reading the whole file first, so the program divides its byte size instead. A boundary will usually land in the middle of a tweet, and two rules make that safe:

  • A rank keeps reading past its chunk_end until its last tweet has a place, so it always finishes what it started.
  • A rank that starts mid-tweet never saw that tweet's _id, so the author and place lines that follow are ignored: its counters are already level.

So every tweet is counted by exactly one rank, with one exception we only found while porting. If a cut lands inside the four spaces of indentation before "_id", the next rank's partial first line still matches the regex, and both ranks count that tweet. At about 2 kB per tweet that is a 1 in 500 chance per boundary on bigTwitter.json. The original Python does it too, so the port keeps the bug and a parity test pins it; the lab flags it when one of your runs hits it.

8
018.7 GB
Rankchunk_startchunk_endSizeTask host
002,341,913,3832.34 GBtask 1
12,341,913,3834,683,826,7662.34 GBtask 2
24,683,826,7667,025,740,1492.34 GBtask 3
37,025,740,1499,367,653,5322.34 GB–
49,367,653,53211,709,566,9152.34 GB–
511,709,566,91514,051,480,2982.34 GB–
614,051,480,29816,393,393,6812.34 GB–
716,393,393,68118,735,307,0602.34 GB–

chunk_size = ceil(file_size / size); the last range absorbs the remainder. Offsets are for the real 18,735,307,060-byte bigTwitter.json.

02 · twitter_processorV1

A line scanner with magic numbers

The file is pretty-printed JSON, one field per line, with a fixed layout. After finding a tweet's _id the scanner skips 2 lines; after author_id it skips 18; after full_name it skips 20. Those counts (commented # MAGICS NUMBERS in the original) jump over text and metadata without even running a regex on them. Below is the real trace of the ported scanner over the first two tweets of a synthetic file.

27 lines regex-tested, 80 skipped unread

  1. 1read:[
  2. 2read: {
  3. 3read: "_id": "1464758956051061163",_id captured 1464758956051061163; skip the next 2 lines
  4. 4skipped: "_rev": "2-5c5d032c9bb4fa870793cd785e77ad94",
  5. 5skipped: "data": {
  6. 6read: "author_id": "589236440400667814",author_id captured 589236440400667814; skip the next 18 lines
  7. 7skipped: "conversation_id": "1464758956051061163",
  8. 8skipped: "created_at": "2021-11-28T00:52:03.684Z",
  9. 9skipped: "geo": {
  10. 10skipped: "place_id": "af7c5ad704ea233b"
  11. 11skipped: },
  12. 12skipped: "lang": "en",
  13. 13skipped: "public_metrics": {
  14. 14skipped: "retweet_count": 0,
  15. 15skipped: "reply_count": 0,
  16. 16skipped: "like_count": 0,
  17. 17skipped: "quote_count": 0
  18. 18skipped: },
  19. 19skipped: "text": "Back home in Newcastle for the first time in ages.",
  20. 20skipped: "sentiment": 0.014230137690902
  21. 21skipped: },
  22. 22skipped: "includes": {
  23. 23skipped: "places": [
  24. 24skipped: {
  25. 25read: "full_name": "Newcastle, New South Wales",full_name → normalised “newcastle nsw”, first sal_dict hit “newcastle” → 1rnsw (Rest of NSW); skip the next 20 lines
  26. 26skipped: "geo": {
  27. 27skipped: "type": "Feature",
  28. 28skipped: "bbox": [
  29. 29skipped: 151.653661,
  30. 30skipped: -33.049339,
  31. 31skipped: 151.898339,
  32. 32skipped: -32.804661
  33. 33skipped: ],
  34. 34skipped: "properties": {}
  35. 35skipped: },
  36. 36skipped: "id": "af7c5ad704ea233b"
  37. 37skipped: }
  38. 38skipped: ]
  39. 39skipped: },
  40. 40skipped: "matching_rules": [
  41. 41skipped: {
  42. 42skipped: "id": 1412189062442586000,
  43. 43skipped: "tag": "synthetic: geotagged tweets from Australia"
  44. 44skipped: }
  45. 45skipped: ]
  46. 46read: },
  47. 47read: {
  48. 48read: "_id": "1516320407265348431",_id captured 1516320407265348431; skip the next 2 lines
  49. 49skipped: "_rev": "1-168b703c45618c8597c915b99d5042f9",
  50. 50skipped: "data": {
  51. 51read: "author_id": "589236440400667814",author_id captured 589236440400667814; skip the next 18 lines
  52. 52skipped: "conversation_id": "1516320407265348431",
  53. 53skipped: "created_at": "2022-04-19T07:38:51.619Z",
  54. 54skipped: "entities": {
  55. 55skipped: "hashtags": [
  56. 56skipped: {
  57. 57skipped: "start": 49,
  58. 58skipped: "end": 54,
  59. 59skipped: "tag": "data"
  60. 60skipped: }
  61. 61skipped: ],
  62. 62skipped: "urls": [
  63. 63skipped: {
  64. 64skipped: "start": 55,
  65. 65skipped: "end": 78,
  66. 66skipped: "url": "https://t.co/054ea03b7e",
  67. 67skipped: "expanded_url": "https://twitter.com/i/web/status/1516320407265348431",
  68. 68skipped: "display_url": "twitter.com/i/web/status/151632..."
  69. 69skipped: }
  70. 70read: ]
  71. 71read: },
  72. 72read: "geo": {
  73. 73read: "place_id": "af7c5ad704ea233b"
  74. 74read: },
  75. 75read: "lang": "en",
  76. 76read: "public_metrics": {
  77. 77read: "retweet_count": 0,
  78. 78read: "reply_count": 0,
  79. 79read: "like_count": 32,
  80. 80read: "quote_count": 1
  81. 81read: },
  82. 82read: "text": "Footy tonight in Newcastle - who else is around? #data https://t.co/054ea03b7e"
  83. 83read: },
  84. 84read: "includes": {
  85. 85read: "places": [
  86. 86read: {
  87. 87read: "full_name": "Newcastle, New South Wales",full_name → normalised “newcastle nsw”, first sal_dict hit “newcastle” → 1rnsw (Rest of NSW); skip the next 20 lines
  88. 88skipped: "geo": {
  89. 89skipped: "type": "Feature",
  90. 90skipped: "bbox": [
  91. 91skipped: 151.662001,
  92. 92skipped: -33.040999,
  93. 93skipped: 151.889999,
  94. 94skipped: -32.813001
  95. 95skipped: ],
  96. 96skipped: "properties": {}
  97. 97skipped: },
  98. 98skipped: "id": "af7c5ad704ea233b"
  99. 99skipped: }
  100. 100skipped: ]
  101. 101skipped: },
  102. 102skipped: "matching_rules": [
  103. 103skipped: {
  104. 104skipped: "id": 1412189062442586000,
  105. 105skipped: "tag": "synthetic: geotagged tweets from Australia"
  106. 106skipped: }
  107. 107skipped: ]

03 · normalise_location + sal_dict

From a place name to a capital city

sal.json maps 15,340 suburb and locality names to Greater Capital City codes such as 2gmel, or rural codes such as 2rvic. The program cleans the keys (brackets, “ - ” and full stops removed), adds every word pair of names longer than two words, and builds a 16,616-key dictionary.

Each tweet's place is lower-cased and normalised, then every combination of its words is tried, shortest first. The first key that exists wins. It is fast and usually right, and the port keeps its quirks: try Macquarie Park, Sydney.

  1. 1

    Lower-case

    “box hill, melbourne”
  2. 2

    normalise_location: drop punctuation, abbreviate state names, squeeze spaces

    “box hill melbourne”
  3. 3

    Try every word combination in itertools order (shortest first) against sal_dict; the first hit wins

    box (no match)hill (no match)melbourne (match)box hillbox melbournehill melbournebox hill melbourne

    3 words → 7 combinations; the first hit was number 3.

  4. 4
    2gmelGreater Melbournecounts towards Tasks 2 and 3

Verified: the original Python resolved this exact place to 2gmel against the full sal.json.

04 · gather_task_tdf + reductions

Gather onto three task ranks

Each rank turns its tweets into three partial tables. Rather than sending everything to rank 0, the program spreads the reductions: rank 0 hosts Task 1, rank 1 Task 2 and rank 2 Task 3. Each host receives its own table first, then one from every other rank in order, and writes one CSV.

Task 1 on rank 0

author_id → tweet count

Sum counts per author, rank with ties sharing the best place (method="min"), keep rank ≤ 10. Ties can therefore produce more than ten rows.

Task 2 on rank 1

gcc → tweet count

Drop unmatched tweets and rural codes (\dr[a-z]{3}), sum per capital city, sort by code. “Other Territories” (9oter) is not rural, so it stays.

Task 3 on rank 2

(author_id, gcc) → tweet count

Count distinct cities per author, order by cities then tweets, take the first ten, and spell out the per-city counts, for example #1879gmel.

Fidelity

Ported, not reinvented

The browser runs a line-for-line TypeScript port of the 2023 code, magic numbers and quirks included. The original Python in coursework/ (analysis logic as submitted) was run outside Spartan (with mpi4py stubbed and the ranks re-enacted in order) to produce reference outputs, and the test suite checks the port against them:

In the test suite (runs in CI)

Fixtures produced by the original Python

  • per-tweet records (id, author, normalised place, gcc) identical
  • per-rank tweet counts identical for 1, 3, 4 and 7 ranks
  • task1.csv, task2.csv, task3.csv and task3_1.csv identical
  • process_salV1 on a sal.json sample: same keys, codes and insertion order
  • every demo place resolves exactly as it does against the full sal.json

Checked locally with the course files

Not committed: course data stays off GitHub

  • the full processed sal_dict: all 16,616 keys and codes identical
  • tinyTwitter.json: records, per-rank counts and all four CSVs identical for 1, 3, 4 and 8 ranks