Skip to content
Tweet Cruncher

COMP90024 · Cluster and Cloud Computing · 2023 Semester 1

Nine million tweets, eight cores, one minute forty-one.

For COMP90024 we wrote an MPI program that split 18.7 GB of geotagged tweets across the cores of the University of Melbourne's Spartan supercomputer, cutting an 11:01 job to 1:41. This site brings it back: the original answers, the scaling story, and the same algorithm running on your own CPU.

Or watch the captioned walkthroughs first
spartan-login · ~/comp90024-a1
$ sbatch slurm/1node8core.bigTwitter.slurm
Submitted batch job 46094406

$ mpiexec -n 8 python main.py -t bigTwitter.json -s sal.json
Wall-clock
00:01:41
Cores
1 node × 8
CPU util.
87.13%
Tweets
9.09M
by 119,439 authors
Input
18.7 GB
one pretty-printed JSON file
Speedup
6.5×
11:01 → 1:41 on 8 cores
Busiest city
Melbourne
2,284,909 tweets

The assignment

Three questions, one very large file

Each student pair had to build a parallel program for Spartan that reads a large Twitter dataset together with a gazetteer of Australian suburbs, answers three questions, and runs on 1 node × 1 core, 1 node × 8 cores and 2 nodes × 4 cores, then explain the timings. Our answers:

What we built

Read bytes, not JSON

Parsing 18.7 GB of JSON on one core is slow, and it does not split well. So the program treats the file as bytes: each rank seeks to its own offset, scans lines with three regular expressions, and skips the fields it does not need using fixed line counts.

Only small count tables travel between ranks at the end, which is why two 4-core nodes ran as fast as one 8-core node.

Walk through the algorithm
  1. 01

    Split

    split_file_into_chunks gives each rank an equal byte range.

  2. 02

    Scan

    twitter_processorV1 reads _id, author_id and full_name, skipping 2, 18 and 20 lines.

  3. 03

    Match

    Place names are normalised and their word n-grams looked up in sal.json.

  4. 04

    Reduce

    Ranks 0, 1 and 2 gather the partial tables and write one CSV each.

Explore

Six ways in

About this project

Coursework, revived

Subject
COMP90024 Cluster and Cloud Computing
University
The University of Melbourne
When
2023 Semester 1, Assignment 1 (social media analytics on Spartan)
Team
Sunchuangyu (Rin) Huang @rNLKJA and Wei Zhao

Academic integrity. The original 2023 submission is kept for reference in the repository's coursework/ folder, with its analysis logic unchanged (it has since been reformatted with black and isort, and a hard-coded email credential was moved to environment variables). The assignment brief, the course datasets and the written report are not reproduced here; the task is paraphrased. If you are taking COMP90024, please do your own work.

Original stack vs revived stack
Area20232026
LanguagePython 3.7TypeScript (strict)
Parallelismmpi4py on Open MPI, SlurmWeb Workers as ranks, page as the interconnect
Datapolars, pandas, NumPyframework-free ports with Vitest parity tests
PlatformSpartan HPC (UniMelb)Next.js 16 static site on Vercel
OutputCSV files and a written reportinteractive charts, map and timelines
Rigourone run per layoutrepeated benchmark with bootstrap intervals, decision records
AInoneoptional, bring-your-own-key, grounded and audit-logged