About the role
Most AI systems in the wild are one hallucination away from an incident and one runaway agent night away from a $40k cloud bill. Someone has to make them not do that.
That's this role. You own the operational side of the AI systems RYB ships into mid-market clients, and the internal tooling (the brain) that makes those systems observable, cost-capped, and safe to leave running unattended. You'll pair with the Fractional CIO Partner on strategy and with the Mid-level AI Engineer on build.
A typical week:
- Eval design and running — every workflow we ship has a real eval harness before it goes live and after every prompt change. You build them, you run them, you set the pass bar
- Observability and cost governance — Langfuse, LangSmith, or similar tracing across every agent call, plus per-workflow token budgets and hard caps so nothing runs away
- Prompt evolution — production-grade prompt versioning, A/B testing, and rollback discipline. Not "vibes-based" changes
- Incident response — when an agent produces something wrong, you're the person who reads the trace, finds the failure mode, and ships the fix
- Client hand-over — teaching a mid-market client's own team how to operate what we've built after we walk away
- Internal R&D on the brain — RYB's mission-control system that runs our own agents, integrations, and monitoring
You're not the person selling the engagement. You're not the person deciding what to build. You're the person the client leans on when the system needs to work reliably for the next three years.
The tech stack
Three platforms cover the bulk of what you'll work with:
- Claude Code — how we ship. Custom Skills, MCP servers, agentic workflows, and internal build velocity. It's the primary agent runtime for both our own brain and most client engagements.
- AWS — the operational hub. Lambda, EventBridge, Secrets Manager, IAM, CDK. Every RYB brain agent and every client-facing integration runs here.
- Microsoft — the client side. Most mid-market AU clients run M365 + SharePoint, and increasingly Copilot. You'll integrate with M365 admin, SharePoint as a RAG source, Azure OpenAI where Anthropic isn't the fit, and Power Automate where it beats Lambda.
Observability sits alongside these — Langfuse, LangSmith, Arize Phoenix, Braintrust. Pick your poison, have a view.
What we're looking for
- 5–8 years in production engineering — TypeScript/Node, Python, or similar. You've been on-call. You've written the runbook.
- Direct production LLM experience — at least one system you've shipped where the LLM was the failure surface you had to design around, not a nice-to-have
- Claude Code fluency — or the confidence you'll be productive in it inside a fortnight. Bonus: shipped custom Skills or MCP servers to production
- AWS depth — Lambda, EventBridge, Secrets Manager, IAM. CDK a strong plus. The RYB brain runs on AWS and stays there
- Microsoft 365 / Azure familiarity — M365 admin, SharePoint, Azure OpenAI, Copilot readiness. You don't have to be an MVP, but you've deployed into a Microsoft tenant more than once
- Practical experience with eval design and observability — you know why "vibes-based" testing doesn't scale
- Cost intuition — you cap workflows in token budgets, tier models by task, and don't ship a reasoning-model loop without knowing what it costs
- Melbourne-based with right to work in Australia
- Bonus: LangGraph, agentic prompt patterns, prior consulting experience
How we work
- Client-first — when there's client work happening, the client's calendar wins
- Boring reliability over cleverness — an agent that runs quietly for eleven months beats one that's brilliant on demo day
- Markdown-first internally; runbooks are markdown; incident reports are markdown
- Honest assessments over polished pitches — read the forms paradox and shadow AI for the operating principles
- Pair on hard things; ship solo on the rest
- Code review is non-optional and not a status game
What we offer
- Day rate genuinely at the top of the AU senior AI engineering band
- Real production responsibility — your systems run in client environments people bill from every day
- Direct line to the fractional CIO on strategy and to the engineering team on build. No PowerPoint hierarchy in between.
- Mix of clients across financial services, professional services, real estate, mid-market manufacturing
- Pipeline currently supports 2–3 days/week ongoing, growing to 3–4 days as more discovery engagements land in FY26-27
What this is not
- Not a research seat. If your favourite question is "how does this model behave on a novel benchmark", RYB isn't the place.
- Not a "prompt engineer" role. Prompt work is a slice of the job, not the whole thing.
- Not a fractional CIO seat — if you want to be in the room owning the client relationship, look at the Fractional CIO Partner role.
How to apply
Apply through this page. Tell us:
- The most production-incident-driven LLM system you've shipped — what went wrong, how you found out, what you changed
- Your current view on the observability stack you'd pick for a mid-market shop starting from zero, and why
- One eval harness you've built and what it caught in the wild
- A short note on the highest per-request cost workflow you've been responsible for and how you brought it down
CV / LinkedIn optional but appreciated. We'll always do a coffee before any commitment.