We are putting language models into products that businesses run on — extracting structured data from unstructured documents, drafting replies inside messaging workflows, and answering questions over a client’s own catalogue and records. That is a different discipline from demos: it has to be right, auditable and cheap enough to run at volume.

What you would work on

  • Retrieval pipelines over client data — chunking, embeddings, vector search, reranking, and the evaluation to prove any of it works
  • Structured extraction from PDFs, invoices, catalogues and scanned documents where the output feeds a real system
  • Tool-using and agentic workflows that take actions in our platforms, with guardrails and human review where the stakes justify it
  • Evaluation harnesses and regression suites, so a prompt change cannot quietly break production
  • Cost and latency work — token budgets, caching, model routing, and knowing when a smaller model is the correct answer

What we are looking for

  • Strong Python, and enough backend engineering to ship a service rather than a notebook
  • Practical experience with an LLM API and at least one retrieval or agent framework, plus a clear view of where those frameworks get in the way
  • Scepticism about your own output. If you cannot measure whether a change helped, you have not finished
  • An instinct for when a language model is the wrong tool and a regular expression would do

Useful, not required

Fine-tuning or adapter training, OCR and document layout models, working with Indian-language text, or experience deploying models under real cost constraints.

Interested?

Send us your CV and a line about why this role. We read every application.

Apply for this role

Other open roles

Tell us what you are building

We will tell you how we would approach it, and whether we are the right fit.

Call us for any enquiry 011 41771877

Start a conversation