AI systems that hold up after the demo ends.

Wanx AI is a research lab that ships. We train and fine-tune ML models, build the data pipelines and agentic systems that run them at real scale, and turn the resulting research into production software — for a fraction of what naive AI infrastructure costs. Remote-first, working with clients globally.

2.5M files indexed in ~3 minutes, millisecond search
4 LLM providers supported — Claude, GPT, Grok, local Ollama
70+ tools across a 5-agent production architecture
99%+ cost reduction versus naive LLM-context routing

What we build

Three capabilities, one team — software, models, and the data layer underneath them.

Enterprise software

Production systems, not prototypes — backend architecture, databases, integrations, and the operational tooling that makes software usable by a real team on day one.

Proof: GarageOS — multi-branch operations platform. FastAPI, PostgreSQL, Next.js, M-Pesa payments built in.

Model training & fine-tuning

When a general-purpose model isn't accurate enough on your data, we adapt one that is — LoRA fine-tuning, evaluation harnesses, and the failure analysis that tells you exactly where a model still breaks.

Proof: R.O.A.D. — Qwen3-VL fine-tuned for historical handwriting recognition on degraded archival records. Hydra II — a production text-to-music model.

Data pipelines & agentic automation

Multi-agent systems that process real volume without the context bill exploding. We design the data layer first, then build agents that only look at what they actually need to.

Proof: AIDAW — 2.5M files indexed in ~3 minutes, 70+ tools across 5 agents, 4 LLM providers.

Selected work

Systems we designed, built, and shipped — with the numbers to back them up.

Data pipelines & agentic automation github.com/ageraustine/daw-agents ↗

AIDAW — Agentic Petabyte-Scale Media Engine

A 5-agent system (Supervisor, Reaper, FFmpeg, Search, Data — 70+ tools) that lets studios process millions of media files through natural language instead of manual pipelines.

  • Observation masking: large datasets stay in SQLite; only metadata reaches the LLM while agents exchange data directly — a 99%+ cost cut versus routing full data through model context.
  • Scale: indexing that handles millions of files (2.5M indexed in ~3 minutes) with millisecond search, across Claude, GPT, Grok, and local Ollama.

GarageOS

A trust-infrastructure platform for multi-branch auto-repair chains — real-time customer transparency, M-Pesa payments, and a weighted Trust Score computed per job.

  • Magic-link repair tracking over WhatsApp/SMS — no app install required.
  • Full job workflow, HR module, and multi-branch HQ visibility on a FastAPI + PostgreSQL + Next.js stack.
Model training & fine-tuning github.com/ageraustine/OCR-ROAD-BARBADOS ↗

R.O.A.D. — Historical HTR Pipeline

A fine-tuning and evaluation pipeline for historical document handwriting recognition, built on Qwen3-VL and tuned for degraded archival records.

  • Asymmetric LoRA — high-rank adapters on the vision tower, lower-rank on the language model.
  • Condition-aware augmentation that scores each document before deciding how to process it.

Who we've built for

Two engagements, two very different problems — the pattern underneath is the same research-to-production pipeline.

The Action Foundation

theactionfoundationkenya.org ↗

A Nairobi-based NGO working toward an Africa where children and youth with disabilities can realise their full potential. We built an early-childhood-development content generation app for them — combining generated background music with unique synthesized narrator voices to produce accessible storytelling content.

Rightsify

rightsify.com ↗

An AI music company building synthetic training datasets and licensing infrastructure for the music industry. We build and maintain Hydra II, their production text-to-music generation model — from research and training through deployment and ongoing evaluation.

How we work

The same four steps on every engagement, in order.

  1. 1

    Scope

    We turn your requirements into a precise technical specification before any code is written.

  2. 2

    Build

    Engineers — not just prompts — implement the system, with evaluation criteria defined upfront.

  3. 3

    Evaluate

    Every model and pipeline ships with a measurable evaluation harness, not a demo that happened to work once.

  4. 4

    Operate

    We hand off with monitoring and documentation in place, or stay on to run it.

Wanx AI is a remote-first research lab founded in Nairobi, working with clients globally. The same rigor we bring to novel research — physics-informed neural networks, neural operators, custom multi-agent RL environments — carries into every client engagement. We're a small team that ships production systems ourselves — no account managers standing between you and the engineers building your system.

Tell us what you're building.

Send a note about the problem you're solving — software, a model, or a pipeline that needs to run at real scale — and we'll get back to you.