COMP90024 · Cluster and Cloud Computing · 2023 Semester 1
Nine million tweets, eight cores, one minute forty-one.
For COMP90024 we wrote an MPI program that split 18.7 GB of geotagged tweets across the cores of the University of Melbourne's Spartan supercomputer, cutting an 11:01 job to 1:41. This site brings it back: the original answers, the scaling story, and the same algorithm running on your own CPU.
Or watch the captioned walkthroughs first$ sbatch slurm/1node8core.bigTwitter.slurm
Submitted batch job 46094406
$ mpiexec -n 8 python main.py -t bigTwitter.json -s sal.json
- Wall-clock
- 00:01:41
- Cores
- 1 node × 8
- CPU util.
- 87.13%
- Tweets
- 9.09M
- by 119,439 authors
- Input
- 18.7 GB
- one pretty-printed JSON file
- Speedup
- 6.5×
- 11:01 → 1:41 on 8 cores
- Busiest city
- Melbourne
- 2,284,909 tweets
The assignment
Three questions, one very large file
Each student pair had to build a parallel program for Spartan that reads a large Twitter dataset together with a gazetteer of Australian suburbs, answers three questions, and runs on 1 node × 1 core, 1 node × 8 cores and 2 nodes × 4 cores, then explain the timings. Our answers:
What we built
Read bytes, not JSON
Parsing 18.7 GB of JSON on one core is slow, and it does not split well. So the program treats the file as bytes: each rank seeks to its own offset, scans lines with three regular expressions, and skips the fields it does not need using fixed line counts.
Only small count tables travel between ranks at the end, which is why two 4-core nodes ran as fast as one 8-core node.
Walk through the algorithm- 01
Split
split_file_into_chunks gives each rank an equal byte range.
- 02
Scan
twitter_processorV1 reads _id, author_id and full_name, skipping 2, 18 and 20 lines.
- 03
Match
Place names are normalised and their word n-grams looked up in sal.json.
- 04
Reduce
Ranks 0, 1 and 2 gather the partial tables and write one CSV each.
Explore
Six ways in
About this project
Coursework, revived
- Subject
- COMP90024 Cluster and Cloud Computing
- University
- The University of Melbourne
- When
- 2023 Semester 1, Assignment 1 (social media analytics on Spartan)
- Team
- Sunchuangyu (Rin) Huang @rNLKJA and Wei Zhao
Academic integrity. The original 2023 submission is kept for reference in the repository's coursework/ folder, with its analysis logic unchanged (it has since been reformatted with black and isort, and a hard-coded email credential was moved to environment variables). The assignment brief, the course datasets and the written report are not reproduced here; the task is paraphrased. If you are taking COMP90024, please do your own work.
| Area | 2023 | 2026 |
|---|---|---|
| Language | Python 3.7 | TypeScript (strict) |
| Parallelism | mpi4py on Open MPI, Slurm | Web Workers as ranks, page as the interconnect |
| Data | polars, pandas, NumPy | framework-free ports with Vitest parity tests |
| Platform | Spartan HPC (UniMelb) | Next.js 16 static site on Vercel |
| Output | CSV files and a written report | interactive charts, map and timelines |
| Rigour | one run per layout | repeated benchmark with bootstrap intervals, decision records |
| AI | none | optional, bring-your-own-key, grounded and audit-logged |