“We'll set up a chatbot for you, and then it just runs by itself.” You hear that a lot. It's wrong. Setting up an AI chatbot is the easy part. Keeping it good for years is the real work. Why ongoing maintenance is not a luxury, and what actually happens behind the scenes.
What happens to an unmaintained chatbot?
Most small-business chatbots are left to their own devices after launch. You can tell because, six months in, they answer worse than they did on day one. Nobody notices, though, because nobody is measuring.
A concrete example from practice: a chatbot correctly answers questions in English at first. After a model update at OpenAI, it suddenly starts answering in English every time, even when the question is asked in German. The bot's operator doesn't notice. The customers don't say anything. They simply go to a competitor.
Silent degradation is the real problem. Nobody complains, everything looks normal. Until one day you notice that enquiries through the chatbot are dropping off. By then the damage is already done.
Here's what a realistic start looks like: we first build the chatbot around your business's 10 to 20 most common questions, things like opening hours, prices, location, contact details, typical services. The bot answers these core questions reliably from day one.
Once it's live, though, questions come in that nobody thought of during setup. At first it answers these worse, sometimes not at all. Averaged across all enquiries, a freshly launched chatbot's quality is therefore usually around 70%. The question isn't whether it starts out perfect. The question is what happens next.
Chatbot quality over 12 months
With vs. without ongoing maintenance. Schematic illustration based on practical experience. Robust independent industry benchmarks for decay or improvement curves do not exist at this level of detail. The trends reflect typical observations from customer projects.
Why do chatbots deteriorate in the first place?
A chatbot isn't a static piece of work. It depends on factors that keep changing:
- Model updates change behaviour. OpenAI, Anthropic and other providers regularly update their AI models. What worked flawlessly yesterday can behave differently today. Sometimes better, sometimes worse, often just different from what you'd expect.
- New questions come up. Customers ask things nobody thought of during setup. Without regularly reviewing conversations, these go unanswered or get answered incorrectly.
- Your content changes. Prices, opening hours, services, staff, locations. The bot has to grow with your business, or it gives out-of-date information.
- Tone drifts. Model updates often change the tone. A bot that used to sound friendly and casual can suddenly come across as stiff and formal after an update. That no longer fits the brand.
- No monitoring, no visibility. Without tools for quality monitoring, you simply don't know what your chatbot is doing. You might see that it's giving answers. Whether those answers are good is a different question.
What we actually do at ServasBot
To keep your chatbot good over the long term, a quality process runs in the background that, ideally, you as a customer never even notice. Specifically:
- Random-sample quality checks. A portion of your chatbot's answers is checked automatically for whether they're factually correct, written in the right language and in the right tone. According to peer-reviewed research (Zheng et al., NeurIPS 2023), such AI-assisted checks reach around 80% agreement with human reviewers, roughly the same level of agreement typically seen between humans. Anything flagged then lands on our manual review list.
- Manual review of flagged items. Whatever the automated check flags, we look at personally. A poor answer gets corrected, the bot's knowledge gets extended, the example gets added to training.
- Model updates tested, not adopted blindly. When OpenAI or another provider releases a new model, we first test it against a gold-standard set of typical enquiries from your business. It only goes live for you once the new model is as good as, or better than, the old one. If it performs worse, the old model stays active.
- Maintaining the knowledge base. As soon as you tell us something has changed (new prices, changed opening hours, new staff), we build it in. The bot always knows the current state of things.
- Monthly quality reporting. Once a month you get an overview: how many enquiries came in, how many were answered well, where there were problems, what we improved.
Self-service vs. managed service: who does what?
There are inexpensive chatbot platforms where you do everything yourself. That works if you have the time and the technical know-how for it. ServasBot's self-service chatbot works on the same principle, from € 49 a month, with a 30-day free trial. This article shows the maintenance effort you should realistically plan for, so you make the decision with your eyes open. Here's the comparison:
| Task | Self-service | ServasBot Managed |
|---|---|---|
| Regularly review answers | You | We do |
| Test model updates | You (or you don't notice) | We do |
| Correct poor answers | You | We do |
| Keep content up to date | You | You tell us, we build it in |
| Quality reporting | Evaluate it yourself | Monthly report |
| Time spent per week | 2 to 4 hours | 0 hours |
Self-service is cheaper in price but more expensive in time. Anyone who takes their chatbot seriously and wants to run it properly needs at least 2 to 4 hours a week for ongoing quality assurance. For a small-business owner with a full-time job on top, that's a real problem.
What you get out of it
- Reliability. Your chatbot stays at the quality level it started at. Ideally it even gets better, because it learns from every customer enquiry.
- No effort on your side. You look after your business, we look after the chatbot. Get in touch if you need to, otherwise it just runs.
- Your customers' trust. A chatbot that consistently gives good answers builds trust. One that answers differently each time damages it.
- No nasty surprises after model updates. You only hear about updates once they reach you, and only when they're an improvement.
What to watch out for
If you're a small business and want to deploy a chatbot, don't just think about what the initial setup costs. Think about who maintains it over the long run, too. A cheap self-service bot that isn't maintained will cost you more customers in the long run than a pricier managed-service bot that keeps working consistently well.
With ServasBot's managed service, ongoing maintenance is a built-in part of the package, no surcharges and no maintenance contract with small print. It's currently at capacity, with onboarding via the waiting list. With self-service you take care of the maintenance yourself, with the time figures from the table above as a realistic guide. Personally supported from Villach, Carinthia, GDPR-compliant with EU hosting.
Data and sources
Solid, independent industry benchmarks for resolution rates of small-business chatbots are rare. Most of the figures circulating come from providers themselves and are not methodologically validated. The following market observations are essentially the reliable ones:
- Gartner (March 2025) predicts that by 2029, agentic AI will autonomously resolve around 80% of common customer-service issues without human intervention. Today, the share is well below that.
- Gartner (December 2024): by 2027, GenAI-powered third-party tools are expected to resolve 40% of customer-service issues.
- Zendesk CX Trends 2026 (a survey of 11,297 consumers and business representatives across 22 countries, June 2025): 74% of consumers expect 24/7 service, and 85% of CX leaders say customers switch providers when issues aren't resolved on first contact.
- Zheng et al. (NeurIPS 2023): Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — the peer-reviewed primary source for the roughly 80% agreement between AI-based quality checks and human reviewers.
The claim that “a well-maintained chatbot can get better over time, while an unmaintained one gets worse” is widely observable in practice, but it isn't quantifiable as a study finding. The chart above should therefore be understood as a sketch based on experience, not as a proven research result.
