1. Why selecting an AI knowledge system is harder than buying one

The market for AI assistants in service has become crowded: every vendor demonstrates a system that answers questions about a machine fluently. The problem is that the demo says almost nothing about whether the system works on your documentation, your service reports and your technicians. The IONOS digitalisation study 2026 (YouGov, around 4,000 decision-makers in five countries) shows what decision-makers care about: for 55 percent of German SMEs, reliability of results is the most important purchase criterion for AI solutions, 43 percent demand compliance with legal requirements, and 36 percent require a provider from Germany or Europe.

Reliability, however, cannot be demonstrated in a demo - only tested. Those who do not test it end up in the statistic Gartner forecast back in 2024: at least 30 percent of generative AI projects will be abandoned after proof of concept - because of poor data quality, inadequate risk controls, escalating costs or unclear business value. None of those four reasons has anything to do with the language model. All four can be tested before you buy.

This checklist grew out of vendor conversations on both sides of the table. It is deliberately written so that you can use it against us, too. If you first want an overview of where AI actually works in service, the guide to AI in after-sales is the better starting point - this article picks up afterwards: you know what you want, and now you have to decide what to buy.

2. Six must-have criteria for the AI assistant in technical service

An AI assistant for field service technicians works under conditions an office chatbot never sees: the user is working on a machine that has stopped, wears gloves, is under time pressure and has no interest in small talk. Those conditions produce six requirements that are not negotiable.

01
Every answer shows its source

Document, page, section - and one click to the original passage. That is not a convenience feature but the precondition for a technician being allowed to trust the answer. A wrong torque value or an overlooked safety release costs more than any search time. Test the opposite case as well: does the system say “I cannot find anything on that” when the information genuinely is not in the sources - or does it invent a plausible answer? In service, the second case is a disqualifier.

02
Drawings, tables and wiring diagrams are content, not decoration

In machinery manufacturing, the decisive knowledge is rarely in the running text. It sits in the exploded view with its item numbers, in the torque table, in the wiring diagram, in the fault code list. Many systems read only the text of a PDF and lose exactly these structures. Test with a question whose answer exists only in a table or a drawing - and with a scanned legacy document of the kind every archive holds.

03
Search across mixed data types

The answer to “fault 4711 on series X at customer Y” is scattered: the fault description in the documentation, the fix in a service report from three years ago, the customer’s variant in the ERP, the spare part in the bill of materials. A usable system connects these sources in the context of the specific machine instead of demanding four separate searches. Ask which source systems can be connected and how the machine context (serial number, variant, retrofits) flows into the answer.

04
Data sovereignty is settled in the contract, not just in the architecture

Service data is competitive capital: it records which machines stand where, what fails and how to fix it. According to IONOS/YouGov, more than half of German SMEs fear their data ending up with AI providers in the USA or China; 53 percent distrust non-European providers. The three questions the contract must answer: Where is the data processed? Who has access? Do your service cases feed a training run that other customers of the vendor benefit from? The third question is the one asked least often.

05
Answer quality is measurable - before and after launch

A system whose quality nobody measures can neither improve nor build trust. The vendor must be able to explain how answer quality is tested on your data before technicians work with it - typically with a set of real questions and known correct answers - and how it is monitored continuously in operation: share of questions answered, share of answers corrected, share of answers the technician actually used.

06
There is a path from correction back into the system

On launch day the system is as good as its sources. After that, the maintenance process decides: How does a new service case get in? How is a wrong answer corrected - and does the correction stick next time? Who in the company is responsible, and how much time does it cost per week? A vendor without concrete answers to these questions is selling a project, not a system.

3. Knock-out criteria in the selection

Not every shortfall weighs the same. Some can be fixed during the project; others mean the conversation can end. The table separates the two:

Observation in demo or proposalAssessment
Answers without a verifiable source referenceKnock-out - unusable in technical service
System answers plausibly even when the information is not in the sourcesKnock-out - wrong answers cost more than no answer
The contract allows your data to be used to train models that other customers also useKnock-out - your service knowledge becomes the vendor’s product
No answer to how answer quality is measuredKnock-out - without measurement, no proof and no improvement
Tables and drawings are not processed, or only as textSerious - clarify whether and when this can be solved in the project
Only one source system can be connected (e.g. documentation only, no service reports)Serious - limits the benefit to a subset of cases
Hosting outside the EU without contractual safeguardsSerious - a data protection and trust issue with customers and technicians
No defined maintenance process, no role named in the companyFixable - but only if you staff the role before the pilot starts
No offline or low-connectivity mode for field useDepends on the site - a must in basements and halls without reception

Assessment from our project practice. Which points are knock-outs for you depends on your deployment scenario - the first four, in our view, apply without exception.

4. Twelve questions for every vendor

These questions fit into a single meeting. They target substance, not features - and the quality of the answers says more about the vendor than any reference list.

4.1 On verifiability

  1. Show me, on one of our documents, how an answer references its source - and what happens when I click on the source.
  2. What does the system answer to a question whose answer is not in our documents?
  3. How does the system handle contradictory sources - say, an old and a revised version of the same instructions?

4.2 On our data

  1. Which of our source systems can you connect - documentation, service reports, ERP, ticketing - and which of those are included in the proposal?
  2. How do you process exploded views, tables, wiring diagrams and scanned legacy documents? Show it on an example from our own archive.
  3. Where is our data processed and stored, who at the vendor and its subcontractors has access, and does our data feed a training run that others benefit from?

4.3 On quality in operation

  1. How do you measure answer quality on our data before the system goes to technicians - and which metrics do we see continuously afterwards?
  2. How does a new service case get into the system, how is a wrong answer corrected, and does the correction stick?
  3. Which role do we need to staff internally, and how many hours per week do you realistically expect?

4.4 On the project

  1. What does the system cost in full in the first year - implementation, data preparation, licences, operation - and which items are variable?
  2. What share of the effort is data preparation, and what happens if our data is in worse shape than assumed?
  3. Under which conditions would you advise us against the project?

The last question is the most important. A vendor who cannot name a condition under which their system does not fit either lacks experience or has no interest in your outcome. Our own answer is in section 6.

5. Setting up the pilot so that it delivers a decision

Most pilots fail not on technology but on their design: they start without a baseline, without acceptance criteria, and with a question space so broad that any result can be read as success and as failure. The sequence that has proven itself consists of four steps, all of which come before the first login.

First: measure the baseline. First-time fix rate, search time per technician per week, share of cases resolved remotely - for two to four weeks, with a tally sheet if need be. Without a baseline, the benefit cannot be demonstrated later. According to Aquant (vendor study, nearly 160 service organisations, over 600,000 work orders), 33 percentage points of first-time fix rate separate the best and the weakest service organisations - 86 versus 53 percent. Where you sit in that range determines how much an assistant can move at all.

Second: build a gold-standard question set. 30 to 50 real questions from the hotline and the field, compiled by your most experienced technicians, each with the known correct answer and its source. At least a third of them should be hard: answers that live in tables, that connect several sources, or that do not exist in the documents at all. This set is your acceptance test - and it is the same set for every vendor you compare.

Third: fix the acceptance criteria in advance. What share of the test questions must be answered correctly and with the right source? How many wrong answers without a warning are acceptable (our answer: zero)? Which metric must move in which direction against the baseline? Whoever sets these after the pilot sets them to fit the result.

Fourth: start narrow. One machine type, one service team, one clearly bounded question space. A narrow pilot with clean measurement beats any enterprise-wide solution on slides. What has proven itself gets extended.

6. When AI support in machinery service is the wrong investment

Honesty belongs in every checklist, so here is our answer to question twelve. In our view there are three constellations in which an AI assistant for service is the wrong investment - or at least the wrong first one:

  1. The documentation is thin or paper-only - and there is no budget to change that. A knowledge system can only answer what is in the data. Digitising and structuring the existing material is then the first step. The Machinery Regulation from January 2027 provides a reason for that anyway - one that cannot be postponed.
  2. Cases rarely repeat. Anyone servicing a handful of special-purpose machines a year whose problems never resemble each other will not recoup the build-up. The benefit scales with the repetition of similar cases across the installed base. A good vendor conversation asks this question among the first.
  3. Nobody in the company can take on the maintenance. A system technicians have distrusted twice is not used a third time. If the owner role cannot be staffed - not even with two or three hours a week - the timing is wrong, not the technology.

A survey by VDMA Software and Digitalisation (February/March 2025, 206 companies) confirms the picture: 45 percent of machinery manufacturers name a lack of staff capacity as a barrier to AI, 44 percent unproven ROI, 42 percent insufficient data quality. All three are reasons that can be settled before the purchase - and none of them gets smaller with a better demo.

7. The scorecard for the vendor comparison

Finally, the checklist in a form you can carry into the vendor comparison. Rate each criterion from 0 (not met) to 3 (demonstrated on your data). A knock-out criterion scored 0 ends the evaluation.

CriterionEvidenceWeight
Source binding of every answer (knock-out)Live on your own document, incl. jump to the original passage3x
Behaviour when information is missing (knock-out)Test question with no answer in the sources3x
Data sovereignty: location, access, no third-party training (knock-out)Draft contract, not presentation3x
Measurability of answer quality (knock-out)Result on the gold-standard question set, operating metrics3x
Drawings, tables, scansTest questions answered only in a table or drawing2x
Mixed source systems and machine contextTest question that connects documentation and service report2x
Maintenance process and internal roleDescribed workflow with hours per week2x
Full first-year costProposal with a line item for data preparation2x
Usability in the field (mobile, low connectivity, gloves)Test on the technicians’ device, not on the projector1x
Willingness to advise againstAnswer to question twelve1x

Weighting from our project practice - adapt it to your deployment scenario. The sum matters less than the “Evidence” column: every score above 1 should rest on something you have seen on your own data.

8. Frequently asked questions

What separates an AI assistant for field service technicians from a chatbot for technical documentation?

A general chatbot answers from the language model. An AI assistant for technical service answers from the company's own documentation, service reports and bills of materials - and shows the source for every answer. Without that source binding, a system is unusable in service, because a plausible wrong answer costs more than no answer.

Which criterion matters most when selecting a system?

Verifiability of every answer: document, page, section, checkable in seconds. According to IONOS/YouGov, reliability of results is the most important purchase criterion for AI solutions for 55 percent of German SMEs. Everything else on the checklist builds on it.

How long should a pilot run?

As long as it takes to test a pre-agreed question set against a previously measured baseline - and no longer. What matters is not the duration but that acceptance criteria, test questions and metrics are fixed before the start. A pilot without acceptance criteria can neither pass nor fail.

Do we need our own model to keep our data safe?

No. What matters is not who owns the language model but where the data is processed, who has access, and whether your service cases feed a training run that others benefit from. Those three points must be settled in the contract - with every vendor, regardless of the technology behind it.

When is an AI assistant in service the wrong investment?

When the documentation is thin or paper-only, when service cases rarely repeat, or when nobody owns the ongoing maintenance. In those cases, structuring the existing material is the first step - and a good vendor conversation ends with exactly that recommendation.

9. Sources

All figures cited in the text with their origin. Studies by vendors with a commercial interest are marked as such. German-language sources are linked in the original - they are the actual evidence.

  • IONOS, digitalisation study 2026 (YouGov, c. 4,000 SME decision-makers with up to 250 employees in five countries, January to March 2026, published 14 April 2026; commissioned by a hosting provider): purchase criteria for AI solutions, distrust of non-European providers. In German.
  • Gartner, press release of 29 July 2024 (analyst firm): forecast on generative AI projects abandoned after proof of concept.
  • Aquant, 2025 Field Service Benchmark Report (vendor study, nearly 160 service organisations, over 600,000 work orders): first-time fix rate of top and bottom performers.
  • VDMA Software and Digitalisation, results of the 2025 AI survey (online survey February/March 2025, n = 206 of 1,960 contacted): barriers to AI adoption. In German.

An AI assistant for field service technicians is not bought in the demo but on the gold-standard question set, the draft contract and the scanned manuals from 2009. Whoever brings those three things to the table before the first vendor sits down no longer needs faith for the decision - only a table. And whoever concludes at the end that the time is not right yet has, in our view, decided just as correctly.