CR-2025-04 · Tested on AWS
Message middleware in C++, from replication to AWS load tests
- Scope
- Leader-follower cluster, queues and topics
- Period
- Apr 2025
- Result
- Built alone in C++17, tested on AWS with distributed load tests
- Status
- Done
- Evidence
- Project
Summary
In April 2025 I designed and built a message-oriented middleware alone, end to end. It's a leader-follower cluster written in C++17, with replication, failure handling and recovery, and I tested it on AWS under distributed load.
Middleware like this sits between other programs, moving messages from whoever sends them to whoever should receive them. Mine handles two patterns. Point-to-point queues send each message to one consumer. Publish-subscribe topics send a copy to every subscriber. So the same cluster works with two different ideas of who a message is for.
If all of that lived on one machine, losing the machine would mean losing the messages. That's why the cluster uses a leader-follower design, for high availability, fault tolerance and data persistence. The leader is the node in charge, and the followers keep copies, so losing one node shouldn't lose the data.
I implemented that behavior myself: the leader and the followers, the replication between them, what happens when a node fails, and how the cluster recovers afterwards.
Between nodes I used gRPC with Protocol Buffers to optimize the communication. Protocol Buffers turn each message into a compact binary format with a fixed schema, and gRPC carries those messages between the nodes. On the outside there's a REST API built with Crow, a C++ web framework, so a client only needs HTTP to talk to the cluster. I managed the project with CMake, and the public repository has the C++17 services, the CMake build files and the generated protobuf and gRPC code.
How I tested it
I tested it on AWS with distributed load testing. Then I went through the system logs to see what the nodes had actually done under that load.
Timeline notes
Four months later, in August 2025, I started at APOLO. I went from writing distributed software to running the clusters that distributed workloads run on.
Timeline
Apr 2025Designed the leader-follower cluster, with point-to-point queues and publish-subscribe topics
- Implemented the leader, the followers, replication, failure handling and recovery in C++17
- Used gRPC and Protocol Buffers between nodes, and added a REST API with Crow
- Tested it on AWS with distributed load tests and went through the system logs
Stack
C++17 · CMake · gRPC · Protocol Buffers · Crow · AWS
Links
Next change review: GRID-EAFIT: energy prediction models, containerized and on AWS Next change review: ASC26 in Wuxi: my part was the infrastructure Next change review: Moving APOLO off VMware without turning anything off