Linux, HPC and AI infrastructure · APOLO · ASC26 · EAFIT

I've run production clusters and the GPU stack for AI.

At APOLO I configured NVIDIA drivers, CUDA and cuDNN on the GPU nodes myself. In one AI project, training on a GPU took 16:02 instead of 54:22.On APOLO's production clusters I scheduled jobs with Slurm and helped compile R, ORCA and ICON. At ASC26 I set up the machines and ROCm drivers. Result: group competition winner.For a year I administered Linux on APOLO 2 and APOLO 3: Ansible to turn raw installs into nodes, FreeIPA for logins, the Proxmox firewall and UFW for access.At APOLO I used Ansible to provision nodes, monitored both clusters with Prometheus and Grafana, and added GitHub Actions checks to our Ansible repos. Docker and AWS come from projects.At APOLO I configured NVIDIA drivers, CUDA and cuDNN on the GPU nodes myself. In projects I moved training to a GPU (54:22 → 16:02) and tuned inference in PyTorch.In projects, I built a leader-follower middleware in C++ and load-tested it on AWS, deployed GRID-EAFIT on EC2 with auto-scaling, and ran Hadoop jobs on EMR.

Clusters
APOLO 2 · APOLO 3
Profile
profile/allprofile/hpcprofile/linuxprofile/devopsprofile/ai-infraprofile/cloud

git checkout profile/pick a role ↓

maincheckout -b proxmox+ 7 machines: ESXi → Proxmox+ kernel and OS tuningmerge → main · downtime 0avg node load 55% → 35%
maincheckout -b profile/hpcASC26 Student Supercomputer Challenge+ ASC26 Student Supercomputer…VMware ESXi to Proxmox at APOLO+ VMware ESXi to Proxmox at APOLOmerge → mainprofile/hpc · deployed
maincheckout -b profile/linuxVMware ESXi to Proxmox at APOLO+ VMware ESXi to Proxmox at APOLOmerge → mainprofile/linux · deployed
maincheckout -b profile/devopsVMware ESXi to Proxmox at APOLO+ VMware ESXi to Proxmox at APOLOA staging environment that mirrors production+ A staging environment that…CI/CD pipelines in GitHub Actions+ CI/CD pipelines in GitHub…merge → mainprofile/devops · deployed
maincheckout -b profile/ai-infraUnifoLM-WMA-0: fine-tuning and faster inference+ UnifoLM-WMA-0: fine-tuning and…AI-driven SDLC orchestration+ AI-driven SDLC orchestrationNeural translation, trained on a GPU+ Neural translation, trained on…merge → mainprofile/ai-infra · deployed
maincheckout -b profile/cloudGRID-EAFIT energy prediction+ GRID-EAFIT energy predictionMovie data with Hadoop MapReduce+ Movie data with Hadoop MapReduceBookstore, from monolith to microservices+ Bookstore, from monolith to…merge → mainprofile/cloud · deployed

Full profileI spent a year in production on APOLO's two clusters, GPU nodes included. The AI side is newer and mostly in projects: GPU training and two multi-agent systems.

01 · whoami

Systems engineer, EAFIT 2026.

A lot of my work sits close to the OS and the network, and my AI projects end up there too.

I'm a systems engineer. I studied at Universidad EAFIT until July 2026, and my last year there overlapped with a year at APOLO, on two production clusters. At EAFIT I was also in two semilleros, student research groups. In the cybersecurity one I exploited and fixed OWASP Top 10 vulnerabilities, and in SCAR, the HPC one, I applied HPC techniques to a Monte Carlo simulation of how Colombia would do at the 2026 World Cup.

A lot of what I've built sits close to the operating system or the network. In projects, that meant a DHCP server written in C on UDP sockets, and a leader-follower middleware in C++. At APOLO, it meant scheduling jobs with Slurm, managing cluster logins with FreeIPA and moving seven machines to Proxmox without downtime.

The AI side came later. In production it's the GPU nodes at APOLO, where I configured the NVIDIA drivers, CUDA and cuDNN myself. The rest lives in projects: I moved a translation model's training onto a GPU, fine-tuned UnifoLM-WMA-0's video-generation model and worked on two multi-agent systems. Even there I keep ending up at the same layer: FP16, a Slurm job that asks for one GPU.

01 · PRODUCTION

0 downtime

I moved seven machines from VMware ESXi to Proxmox at APOLO.

02 · GPU

CUDA on the nodes

I configured the NVIDIA drivers, CUDA and cuDNN on APOLO's GPU nodes myself.

03 · ASC26

Wuxi, May 2026

Group competition winner. I set up the infrastructure.

02 · git log --career↕ sorted for HPC↕ sorted for Linux admin↕ sorted for DevOps / SRE↕ sorted for MLOps / AI↕ sorted for Cloud

Two clusters and one competition.

  1. release/2026.07 · HEAD

    Freelance

    Since Jul 2026

    Independent developer now Freelance

    I build custom software for clients, three so far, and every change goes through GitHub Actions and a staging environment before production.

  2. release/2026.05 · tag asc26

    ASC26 Student Supercomputer Challenge

    ASC26 · Global Finals · May 2026

    I set up the machines. The team tuned ICON. Award

    In Wuxi I was HPC infrastructure lead and installed the ROCm GPU drivers. With teams from four international universities, we optimized ICON's inference time and were the group competition winner.

    Where
    Wuxi, China
    My role
    HPC infrastructure lead: I set up the machines and installed the ROCm drivers
    Result
    group competition winner
    Team benchmark
    HPL reached about 78% of Rpeak
    Read the change review
  3. release/2025.08 → 2026.07

    Scientific Computing Center APOLO

    Aug 2025 – Feb 2026 · Supercomputing Analyst
    Feb 2026 – Jul 2026 · Supercomputing Research Assistant

    A year on two production clusters Production

    • Migrated seven machines from VMware ESXi to Proxmox with zero downtime. Hypervisor licensing went to $0.
    • Tuned kernel and OS parameters. On equivalent workloads, average node load went from 55% to 35%.
    • Ran Slurm for job scheduling and FreeIPA for identity management on APOLO 2 and APOLO 3.
    • Used Ansible to turn raw OS installs into working cluster nodes, and to deploy Prometheus.
    • Configured the Proxmox firewall and UFW, and kept the VM backups on the NAS.
    • Monitored both clusters with Prometheus and Grafana.
    • Helped compile and integrate R, ORCA, ICON and ML workloads on the clusters.
    • Added GitHub Actions checks to private repos so new Ansible code followed our rules, including linting and no exposed secrets.
    • Gave induction talks and mentorship on cluster use to ~20 people.
    • Configured the NVIDIA drivers, CUDA and cuDNN on the GPU nodes myself.
    • Set up Lmod and Conda environments for users, and administered Slurm partitions and nodes, including drain and resume.
    • Administered Rocky Linux on the nodes, configured the VLANs and switches, ran the HPL and HPCG benchmarks myself and handled the help-desk tickets.
    • Configured VM templates with cloud-init myself, added Prometheus exporters (node_exporter and a Slurm exporter), and set up systemd services or cron jobs, depending on the need.
    • Worked with Intel MPI, OpenMPI, OpenMP, the Intel oneAPI compilers and Apptainer containers, and helped researchers use them.
    • Helped users with Slurm's GPU scheduling (GRES), QoS and fairshare, the NFS shared filesystem, and remote access over VPN and SSH keys.
    • Supported cluster users from other countries too, such as Poland and Australia.
    • Helped with racking and cabling nodes, Warewulf network installs, InfiniBand, the cluster's internal DNS and DHCP, and user accounts and quotas. Someone else led that work.
    • Slurm
    • FreeIPA
    • Ansible
    • Proxmox
    • Prometheus / Grafana
    • CUDA
    • Rocky Linux

02a · git diff --stat

The numbers and where they came from.

  • 0

    downtime moving 7 machines from VMware ESXi to Proxmox

  • 55% → 35%

    average node load, equivalent workloads, measured in production at APOLO

  • 54:22 → 16:02

    training time without and with a GPU, in an AI project

  • ~78%

    of Rpeak on HPL at the ASC26 Global Finals, a team result

  • ~20

    people got induction talks and mentorship from me on cluster use

03 · ls projects/↕ sorted for HPC↕ sorted for Linux admin↕ sorted for DevOps / SRE↕ sorted for MLOps / AI↕ sorted for Cloud

Things I've built.

The APOLO migration ran in production. Everything else is a project, a course assignment or a competition, and each card says which. Four have a full change review.

  • Production

    VMware ESXi to Proxmox at APOLO

    At APOLO I migrated seven machines from VMware ESXi to Proxmox with zero downtime, and hypervisor licensing went to $0. I also tuned kernel and OS parameters. On equivalent workloads, average node load dropped from 55% to 35%. For access control I set up the Proxmox firewall and UFW, and the VM backups went to the NAS.

    Proxmox · VMware ESXi · Linux kernel tuning · UFW · NAS

  • Award

    ASC26 Student Supercomputer Challenge

    In May 2026 I went to the ASC26 Global Finals in Wuxi, China, as HPC infrastructure lead. My part was the competition infrastructure: I set up the machines and installed the ROCm GPU drivers. The team optimized inference time for ICON, a weather and climate application, working with teams from four international universities. On HPL, the team reached about 78% of Rpeak. We were the group competition winner.

    ROCm · AMD GPUs · ICON

  • Project

    High-availability message middleware

    I designed and built this one alone, end to end, in C++17 with CMake. Nodes form a leader-follower cluster for high availability, fault tolerance and data persistence, and they talk over gRPC and Protocol Buffers. Clients get point-to-point queues and publish-subscribe topics through a REST API written with Crow. I tested it on AWS with distributed load tests and went through the system logs.

    C++17 · CMake · gRPC · Protocol Buffers · Crow · AWS

  • Project

    GRID-EAFIT energy prediction

    A team project at EAFIT that predicts electricity consumption and solar production. My part was the Flask backend, the Keras and TensorFlow models, the Docker Compose setup and the AWS deployment. It ran on EC2 with routing, load balancing and auto-scaling, without Kubernetes.

    Flask · React · Keras · TensorFlow · Docker Compose · Pytest · AWS EC2

  • Project

    A DHCP server from scratch

    I built a DHCP server from scratch in C, on UDP sockets. It hands out IP addresses and manages leases, and it runs on Linux and in simulated NAT environments. No DHCP library does the work: packet handling and network-interface management go through low-level system calls, which is the part of networking most tools hide.

    C · UDP sockets · Linux · NAT

  • Project

    Neural translation, trained on a GPU

    An English-to-Spanish translator in TensorFlow and Keras: an encoder-decoder with LSTM and attention, on top of pretrained spaCy embeddings. Before the GPU, training took 54:22. I moved it to a GPU with CUDA and cuDNN, through a Slurm job that requests one GPU, and training time dropped to 16:02.

    TensorFlow · Keras · spaCy · CUDA · cuDNN · Slurm

  • Project

    UnifoLM-WMA-0: fine-tuning and faster inference

    UnifoLM-WMA-0 is a public world-model framework for general-purpose robot learning. I fine-tuned its video-generation model on the Open-X dataset, and also worked with five Unitree Robotics datasets. On the performance side, I configured FP16 precision and changed how PyTorch's scaled dot-product attention (SDPA) behaved, to cut inference time.

    PyTorch · FP16 · SDPA · Open-X

  • Project

    CI/CD pipelines in GitHub Actions

    Two GitHub Actions setups on Ubuntu with Python 3.12. The larger pipeline is a fork, so I started from someone else's workflow. In it, Black, Pylint and Flake8 check the style, Pytest runs unit and acceptance tests with coverage, SonarCloud reports on code quality, and Docker Buildx with QEMU builds the image that gets published to Docker Hub. The smaller one handles push triggers, manual dispatch and script runs.

    GitHub Actions · Pytest · SonarCloud · Gunicorn · Docker Buildx · QEMU · Docker Hub

  • Project

    AI-driven SDLC orchestration

    I worked on a multi-agent system built with Claude Code that takes a business brief through the software lifecycle and ends with a deployed Django application. Each stage has its own specialized agent. At the validation checkpoints a person has to approve the output before it goes to the next agent.

    Claude Code · Django

  • Project

    A staging environment that mirrors production

    In a university project I built a staging environment inside a container that mirrors production, so changes get tested under realistic conditions before they reach the real thing. It's still running today. I work the same way elsewhere: GitHub Actions validates the code, and staging comes before production.

    Containers · Staging

  • Project

    Movie data with Hadoop MapReduce

    Four MRJob programs in Python that compute director counts, director ratings, actor counts and letter frequencies over a movie dataset. I ran them as Hadoop MapReduce jobs over HDFS on AWS EMR, with the setup done in shell, and put a small Flask interface on top for the results.

    Python · MRJob · Hadoop · HDFS · AWS EMR · Flask

  • Project

    AgentSprint by ReshapeX

    A team challenge on multi-agent systems, where we placed 5th. Our solution takes a natural-language description or a video of a home and produces a 3D simulation that recommends where Xiaomi sensors and smart-home products should go.

    Multi-agent systems · Multimodal input · 3D simulation

  • Coursework

    Bookstore, from monolith to microservices

    A two-person course project. We split a Flask and SQLAlchemy bookstore monolith into three services (authentication, catalog and transactions), each with its own Dockerfile, run together with Docker Compose on AWS EC2. The services still share one Amazon RDS database, and their URLs are hardcoded. We wrote both trade-offs down in the documentation.

    Flask · SQLAlchemy · Docker Compose · AWS EC2 · Amazon RDS

  • Project

    Gold price direction with XGBoost

    A two-day personal project that ran as a Slurm batch job on APOLO's cluster, outside my APOLO work. The first version, an XGBoost regressor on 15-minute gold candles, used a random split on time-series data, so its scores don't count. The second predicts direction over the next four hours, with about 50 technical features and a chronological split. It has no recorded results.

    Python · pandas · XGBoost · scikit-learn · Slurm

  • Project

    Real-time facial emotion recognition

    I trained a CNN with TensorFlow and Keras to classify seven facial expressions, then connected it to a webcam app: OpenCV handles the camera, and the interface is Tkinter with CustomTkinter. The README reports 79% validation accuracy on the facial-expression dataset.

    TensorFlow · Keras · OpenCV · Tkinter

  • Coursework

    Operating systems course projects

    Three projects from my operating systems course. One simulates continuous memory allocation with a Worst Fit policy, covered by Pytest. The page-replacement midterm (procesar, LRU eviction) started from the instructor's template. The last is pysync, a thread-synchronization library with producer-consumer and rendezvous patterns, tested with Pytest and parameterized.

    Python · Pytest · threads

  • Coursework

    Numerical methods and graphs

    Two course projects. SolverPro is a web app in React and TypeScript, built with Vite, with methods for root finding, linear systems, interpolation and differential equations. The other one builds a graph from street data in Python and finds routes with Dijkstra, with pandas for the data and gmplot for the map.

    React · TypeScript · Vite · Python · pandas · gmplot

04 · cat stack.yaml↕ sorted for HPC↕ sorted for Linux admin↕ sorted for DevOps / SRE↕ sorted for MLOps / AI↕ sorted for Cloud

What I've used, and where.


By evidence level

I used the production tools on APOLO's clusters. The project tools come from university and personal repos, and I've only used the coursework ones in classes.

stack.yaml · production · project · courseworkprofile: allhpclinuxdevopsai-infracloud
stack:  production:    - Linux  # Admin on APOLO 2 and APOLO 3    - HPC cluster operations  # Two production clusters at APOLO    - Slurm  # Job scheduling on APOLO 2 and 3    - Proxmox / VMware ESXi  # 7 machines migrated, zero downtime    - Ansible  # Node provisioning, Prometheus deploy    - FreeIPA  # Identity management on APOLO    - Proxmox firewall / UFW  # Hardening on APOLO    - Prometheus / Grafana  # Cluster monitoring on APOLO    - GitHub Actions  # Lint and secret checks for Ansible    - R / ORCA / ICON  # Helped compile and integrate them    - NVIDIA driver / CUDA / cuDNN  # Configured on APOLO's GPU nodes    - VM templates / cloud-init  # Configured myself at APOLO    - node_exporter / Slurm exporter  # Prometheus metrics on APOLO    - systemd / cron  # Services or scheduled jobs, by need  project:    - C / C++  # DHCP server in C, HA middleware in C++    - gRPC / Protocol Buffers  # Node traffic in my HA middleware    - AWS (EC2, EMR)  # GRID-EAFIT on EC2, Hadoop jobs on EMR    - Keras / TensorFlow  # Energy models, translation, emotion CNN    - Docker / Docker Compose  # GRID-EAFIT and Bookstore containers    - Python / Bash  # Hadoop jobs, Flask backends, EMR setup    - AWS auto scaling, Lambda, IAM, S3, ECS, ECR  # Set up myself in project work    - Let's Encrypt  # TLS certificates I set up in projects    - Nginx  # Project work, not production  coursework:    - Terraform  # Coursework, not production    - Kubernetes  # Courses + small EKS test clusters    - Rust  # A forked guided course, still learning profile: allhpclinuxdevopsai-infracloud

05 · ls -la credentials/

Degree, research groups and certificates.

Achievements

  • 2026

    ASC26 · Wuxi

    Team competition Award

    ASC26 Global Finals in Wuxi: group competition winner. I set up the infrastructure.

  • 2026

    AgentSprint by ReshapeX

    Team competition Award

    AgentSprint by ReshapeX: 5th place with my team, for a multi-agent smart-home system.

  • 2022

    Generación E scholarship

    Scholarship Award

    Generación E scholarship recipient.

Education

  • Systems Engineer

    Universidad EAFIT

    Coursework in systems programming, networks, distributed systems, cloud, containers, machine learning, cybersecurity, Linux administration and HPC.

    Jan 2022 – Jul 2026
  • SCAR, Semillero de Computación de Alto Rendimiento (HPC research group)Training

    Universidad EAFIT

    I applied HPC techniques to a Monte Carlo simulation of how Colombia would do at the 2026 FIFA World Cup.

    Aug 2025 – Jul 2026
  • Semillero de Ciberseguridad (cybersecurity research group)Training

    Universidad EAFIT

    Blue, red and white team exercises. I exploited and fixed OWASP Top 10 vulnerabilities.

    Jan – Dec 2025

06 · ping mauricio

Email is the fastest way to reach me.

github: MauricioCa07 · linkedin: mauricio-carrillo-0452a5291