about
I build AI that people can actually rely on
I am an AI engineer with five years of industry experience. I build and evaluate AI systems: multi-agent orchestration, retrieval, evaluation and the guardrails that make them safe to put in front of real users. My deepest specialisation is finance and accounting, because I worked in it before I automated it.
From the books
to the systems

Getting a model to produce something impressive is the easy part now, and it has been for a while. The hard part is everything around it: knowing whether the output is right, catching it when it stops being right, deciding what the system is allowed to do on its own, and giving a person a way to see what happened and disagree. That is where I spend almost all of my time, and it is why my work is organised around reliability rather than capability.
I learned that lesson somewhere specific. Before software I worked in accounting, reconciling accounts and closing the month, so I have been the person who finds the mistake three weeks later when the numbers are already signed. It taught me what an unreliable system actually costs, which turns out to generalise well beyond accounting. If you want that thread in full, it is on the finance page.
I came up through full-stack TypeScript and React before Python and LangGraph, which turns out to matter more than it sounds. An AI feature is not finished when the agent works; it is finished when a reviewer can see what the agent did and disagree with it. Being able to carry a feature from the graph through the API and into the screen is what makes that possible rather than aspirational.
Alongside the engineering I lead a faculty of around a hundred instructors at the Governor Sindh Initiative and teach at PIAIC and Saylani. Thousands of engineers have come through those programmes. Teaching keeps me honest: you cannot hand-wave a concept to a room of three hundred people who are about to try it themselves.
How I work
I use written specifications, evaluation suites, and browser checks to test engineering decisions. Each case study explains the alternatives and the evidence behind the choice. Those details are part of the work, not something added after it.
Why everything here is synthetic
Every project on this site runs on generated books. Companies, counterparties, invoices and bank transactions produced from a seed by a data engine I wrote for exactly this purpose. None of it corresponds to a real organisation, and nothing is ever posted to a real system.
That is a deliberate constraint, not a limitation I worked around. Client financial data stays with clients; publishing a portfolio built on it would be a breach whatever the engineering merit. The alternative most people take is to demo on data so clean it proves nothing, so the generator deliberately produces books that are wrong in the ways real books are wrong: mixed date formats, missing references, aliased payee names, duplicate payments, and payments that nothing in the ledger explains.
The honest cost is that a perfect score on synthetic data is a much weaker claim than a good score on real books, and every case study says so in its own words rather than leaving you to infer it.
Where I am now
AI Engineer at Voya AI, remote from Karachi, Pakistan, building multi-agent systems proven against accounting and finance work. Open to fully remote roles worldwide. Agentic AI, applied AI and LLM teams, in any domain where agents must be trusted. If you are drowning in month-end work, or trying to work out whether the agents you already have can be trusted, that is the conversation I want to have.