Pick a real business job.
“Use AI” is not a brief. “Turn every sales call into a clean summary, follow-up draft and CRM update within two minutes” is.
Implementing AI / After the hype
The buzzword has been beaten to death by people who automated one email and called themselves futurists. Underneath the noise is genuinely useful infrastructure. Let us show you the sauce.
Skip the keynote ↓First principles
A model can read, reason, classify, draft, compare, extract and sometimes use tools. That does not mean it understands your business, should touch every process or gets permission to freestyle inside the bank account.
Useful implementation starts with a real job, gives the model the right context, picks the right level of intelligence, limits what it can do and measures whether the workflow is better than before.
01 / The actual sauce
The chat box is the visible bit. The quality comes from the system around it.
“Use AI” is not a brief. “Turn every sales call into a clean summary, follow-up draft and CRM update within two minutes” is.
The model needs the right slice of your offers, customers, tone, SOPs, current records, examples and constraints at the moment it works. More context is not always better. Relevant context is.
State the outcome, inputs, constraints, process, examples and definition of done. Prompting is not collecting magic phrases. It is learning to specify work clearly and test the output.
Do not pay frontier-model prices to label a support ticket. Do not ask the cheapest tiny model to make a high-stakes commercial judgement. Different work deserves different brains.
Reading a task list is different from changing it. Drafting an action is different from executing it. Expose only the tools required, keep an audit trail and put approval in front of risky writes.
Track accuracy, time saved, correction rate, latency, token cost and commercial outcome. If humans spend longer fixing the output than doing the work, the demo failed.
02 / Things we have actually built
Neo HQ + Hermes
We built a private operations HQ, then connected a Hermes agent running on a VPS through a controlled tool bridge. It could discover authorised tools for tasks, clients, pipeline, revenue context and ad-account status without opening the entire system to the internet.
Operator HQ approvals
For an operating dashboard, we created an agent bridge that returned normal advice plus structured proposed actions. The first executable write surface was deliberately narrow: create a task only after explicit approval.
03 / What this could look like for you
Say you run a service business and leads arrive through forms, phone calls and email. Here is a useful implementation. Not a chatbot waving from the bottom-right corner.
The system validates the contact details and pulls the source, page, service and campaign context.
Job type, location, urgency, likely value and missing information are returned in a strict structure.
High-intent work goes to the right person. A useful reply is drafted in your tone. The CRM record and follow-up task are prepared.
A simple acknowledgement can send automatically. A quote, promise or unusual case waits for a human.
Booked, lost, junk or no-answer data returns to the system so prompts, routing and acquisition improve.
04 / Different models, different work
Frontier reasoning
Strategy, complex planning, messy document analysis, coding and decisions where a better answer is worth the extra latency and cost.
Fast, lower-cost models
Classification, extraction, formatting, routing, first drafts and repetitive work with clear rules and cheap verification.
Multimodal & specialist models
Calls, screenshots, plans, invoices, photos, video or realtime voice. The best text model is not automatically the best model for every medium.
Open & Chinese models
Qwen, DeepSeek, Kimi and open-weight models can be strong, cost-effective parts of a router. They are not automatically the right choice. Data handling, hosting, reliability, provider risk and output quality still need testing.
Model loyalty is not a business strategy. Route by quality, speed, privacy and cost—then keep an exit door.
Use cheaper models for clear, high-volume work.
Send the relevant records, not the entire company drive.
Reuse stable instructions and repeated context where providers support it.
Run non-urgent work efficiently instead of demanding instant answers.
05 / Context engineering
A good prompt with bad context produces polished guessing. The goal is to assemble the smallest trustworthy packet of information needed for this job, right now.
Offers, margins, customers, territory, policies, tone and what “good” actually means.
The current lead, job, client, account, task, document or conversation.
Outcome, constraints, examples, tools, edge cases and a definition of done.
Permissions, approvals, structured output, logging, evaluation and a fallback path.
06 / Prompting without wizard cosplay
One outcome. One owner. One definition of done.
Real examples beat ten paragraphs of adjectives.
What it may use, must avoid and should do when uncertain.
Use fields, schemas or checklists when another system consumes the result.
Missing data, contradictions, nonsense input and customers doing weird customer shit.
Prompts are operating logic. Change them deliberately and compare outcomes.
07 / Find a useful first move
Tell us what repeats, what requires judgement, where the data lives and what going wrong would cost. We will map a sensible first implementation—even if the answer is “do not use AI here.”
hello@goodproblems.com.au