AI and LLM application engineering
Most assistants are a prompt with a text box in front of it. The ones worth paying for are grounded in something — a user's actual history, a document set, the state of the system they are talking about. That grounding, and the plumbing that keeps it affordable and honest, is the work.
- Retrieval and grounding: an assistant that opens with what the user actually got wrong, rather than a generic first message
- Model routing across providers, with key rotation, health-based retirement and automatic failover, so an outage slows a job rather than losing it
- Credit and quota metering with a ledger that reconciles, prices shown before the action, and automatic refunds on failure
- Support for hosted and local models through OpenAI-compatible endpoints, including fully air-gapped deployment
- Robust handling of malformed model output, and checkpointing so a long job resumes rather than restarts