◆ Available for freelance & consulting

Work with me

I build AI systems that have to be right — and prove they are. If you have an LLM feature that works in a demo but not in production, an integration nobody wants to maintain, or a bill that grew faster than the usage, that's the work I do.

Everything I offer below is something you can already see working on this site. No capability listed here is one I can't point at.

What I do

Five things, each with a piece of work behind it

Agentic systems & workflow automation

Multi-step agents that route, decide and escalate — with a human in the loop wherever a wrong answer would be expensive.

  • Router and orchestration design across your existing channels and tools
  • Human-in-the-loop approval paths for anything high-stakes
  • Observability so you can see what the agent did and why
Proof: Omni-channel assistant

RAG & document AI pipelines

Turning messy documents into structured, validated data your systems can actually rely on.

  • Ingestion and layout-aware parsing for real-world file formats
  • Strict schema extraction with validation at the boundary
  • Deterministic verification — grounding, plausibility bounds, self-consistency
Proof: Document AI experience

MCP servers & integrations

One standard interface between your internal systems and any AI client, instead of bespoke glue for every tool.

  • MCP server exposing your data and actions safely
  • Tool design that an agent can actually use without hand-holding
  • Deployment on your infrastructure, with the access boundaries written down
Proof: A live MCP server you can connect to

LLM evaluation & reliability

Knowing whether your AI feature works — measured, not assumed — before your users find out for you.

  • Golden dataset and rubric design for your domain
  • Evaluation harness that runs on every change, with the numbers tracked
  • LLM-as-a-Judge scoring validated against human review
Proof: LLM-as-a-Judge project

Cost & latency engineering

Same output, materially smaller bill — for teams already running LLMs in production and feeling it.

  • Token accounting to find where spend actually goes
  • Prompt caching, context compaction and model-tier routing
  • Local / open-weight model options where they hold up
Proof: Local-first PII redaction

How we'd work

Pick the shape that fits — scope and price agreed before anything starts

Technical review

Fixed price · about a week

You have something built or half-built and want an honest read: what will break, what it will cost at volume, what to do next. Written findings you can act on.

Proof of concept

Fixed price · two to four weeks

One narrow, real use case taken end to end so you can judge feasibility on evidence rather than a demo video.

Build & hand over

Scoped project

A production pipeline with evaluation, documentation and a handover session, so your team owns it after I leave.

Ongoing consulting

Day rate

Regular time for architecture review, pairing and unblocking while your team builds the thing themselves.

  1. 01

    Call

    Half an hour on what you're trying to do and whether I'm the right person. No charge.

  2. 02

    Written scope

    What I'll deliver, what it costs, how long, and what I need from you. Agreed before anything starts.

  3. 03

    Build in the open

    Regular check-ins and working code you can see, not a reveal at the end.

  4. 04

    Handover

    Documentation, evaluation results and a walkthrough — so the work outlives the engagement.

Tell me about it

A couple of sentences is plenty — I reply to everything within a couple of days

or email zaahidimraan@gmail.com

Your details are stored only so I can reply. No newsletter, no third-party analytics, nothing passed on.

Questions

The ones I get asked most

What AI engineering services do you offer?

Agentic system design and build, retrieval-augmented generation (RAG) pipelines, Model Context Protocol (MCP) servers and integrations, LLM evaluation harnesses, and cost optimisation for teams already running LLMs in production.

Do you work with startups as well as established companies?

Yes. Engagements range from a short technical review or proof of concept for a small team, through to building and handing over a production pipeline with an evaluation suite.

How do you price freelance AI engineering work?

Fixed price for scoped pieces of work such as a proof of concept, an evaluation harness or an MCP integration, and a day rate for ongoing consulting. Scope and price are agreed before any work starts.

What is an MCP server and why would my team want one?

A Model Context Protocol server exposes your internal systems to AI clients through one standard interface, so an assistant can query your data without bespoke integration code for every tool. This portfolio publishes one you can connect to and try.

How do you make sure an LLM feature is actually reliable?

By measuring it. That means a golden dataset, an evaluation harness that runs on every change, deterministic verification around model output where correctness matters, and grounding checks so answers can be traced to a source.

Where are you based and do you work remotely?

Manchester, United Kingdom. Remote work across UK and European time zones, with on-site available in the North West.