AI in replies
AI guardrails: stop the bot promising what you don't have

An AI assistant invents details not because it’s «broken» but because it was handed a question without data and never told that «I don’t know» is an acceptable answer. Guardrails are the explicit list of prohibitions and stop-words that turns a guess into a handover.
One invented promise costs more than a dozen missed messages: the customer arrives with a screenshot, and there is nothing left to argue about.
Which promises should AI never make?
Four categories where an apology doesn’t fix the damage:
- 1
Timelines
«It'll be there Friday» without checking stock and courier. The assistant quotes only dates that exist in data.
- 2
Prices and discounts
Any figure not in the catalogue — including «we'll do it cheaper» in response to haggling.
- 3
Specifications
Materials, compatibility, warranty, origin. What isn't in the description doesn't exist for the assistant.
- 4
Availability
«Yes, we have it» without real inventory access. The most common and most expensive invention.
How do you phrase a guardrail so it actually holds?
A prohibition works when it can be checked unambiguously. «Be careful with prices» is not an instruction. «Quote only prices from the catalogue price field; if the field is missing, say you’ll check and hand the conversation to a manager» is.
The test: could an outsider read the rule and say yes or no about a specific reply? If not, rewrite it.
What are stop-words and why do you need them?
Words after which the assistant doesn’t answer at all — it calls a human. Not «answers more carefully». Goes quiet.
A baseline set for a shop: complaint, faulty, defective, refund, scam, lawyer, court, police, plus anything involving health, children and safety. Every niche adds its own: cosmetics — allergy; electronics — smoking, sparked; food — poisoning.
The logic is simple. In these conversations the cost of a mistake isn’t a lost sale but a public dispute or a legal problem. Automation here doesn’t save money, it manufactures risk.
Can AI message first?
Only inside the limits the platform itself sets, and this isn’t a matter of taste. On Messenger a business replies within 24 hours of the customer’s message; beyond that window there is a separate extension of up to 7 days, and it is permitted only for messages written by a human agent — automated messages are explicitly not allowed there (Meta, Messenger Platform policy, 2026).
So «nudging» a customer with a bot on day four isn’t aggressive marketing, it’s a platform-policy breach that puts your page at risk. Follow-up outside the window is written by a person.
Minimum guardrail checklist before launch
- Data source: the assistant answers only from the catalogue; outside it — «I’ll check».
- Forbidden promises: timelines, prices, discounts, specs, availability are never named without data.
- Stop-words: the list that sends a conversation to a human with no attempt to answer.
- Attempt limit: three repeated questions in one thread — handover.
- Tone: no apologising for being automated, no manufactured empathy, no «I’ll personally make sure».
- Outside Messenger’s 24-hour window, no automated messages are sent.
Frequently asked questions
What are AI guardrails in plain terms?
An explicit list of what the assistant never does: never quotes a price that isn't in the catalogue, never promises a delivery date, never agrees to a discount, never invents a spec. Anything outside the allowed list goes to a human.
Why does AI invent product details?
Because it was asked to answer without being given data, so it fills the gap with the most plausible text. The fix isn't a «better model» — it's catalogue access plus explicit permission to say «I'll check and get back to you».
Can you let AI give discounts?
Only inside hard numeric limits you set — say a fixed 5% above a certain order value. Open haggling belongs to a human: the assistant can't judge when a concession is earned and when you just gave away margin.
Which words should trigger immediate handover?
Complaint, faulty, refund, lawyer, court, plus anything touching health and safety. Add your own for your niche. These are stop-words: the assistant goes quiet and calls a person instead of attempting an answer.
How do you test that guardrails hold?
Once a week, deliberately ask the assistant about something that doesn't exist: a fake size, a 50% discount, a ten-year warranty. Correct behaviour is admitting it doesn't know and handing over. If it improvises, your rules aren't specific enough.



