Support teams are training AI on data even data leaders don't trust yet
A joint report from Snowflake and Omdia, covered by No Jitter on April 1, 2026, found that 79% of organizations face data-centric challenges when building AI systems — yet 92% are already using that same data to train large language models. The most commonly cited problems were breaking down AI data silos (65%) and measuring or monitoring AI data quality (62%). Separately, 40% of respondents named data quality their top concern. Snowflake's VP of AI, Baris Gultekin, framed the gap plainly: companies "are not waiting until everything is perfectly clean and ready, because they can't afford to."
What did the survey actually find?
Most organizations know their AI data isn't ready, but are training and deploying models on it anyway rather than waiting to fix it.
The 79%-vs-92% gap is the whole story: readiness concerns and actual deployment are moving in opposite directions. Silos (65%) and quality monitoring (62%) aren't fringe complaints — they're the two things a support AI depends on most, since a chatbot or agent-assist tool is only ever as good as the knowledge base, ticket history, and CRM fields it's pulled from.
Why are teams shipping anyway?
Competitive pressure and vendor roadmaps are outrunning internal data hygiene, and few teams see a realistic path to "clean enough" before the next release cycle.
That's consistent with what this site has already flagged about AI customer service more broadly: only a fraction of deployed use cases are producing measurable ROI. If nine in ten companies are training on data they themselves rate as siloed or hard to quality-check, weak ROI numbers stop looking like a rollout problem and start looking like a data problem that got shipped anyway.
What does this mean for a support operation specifically?
Answers your AI gives customers are drawing on the same unresolved silos and quality gaps IT teams admit they haven't fixed.
Support data is scattered by design — help center articles, macros, past tickets, CRM notes, product docs — often owned by different teams with different update cadences. If 65% of organizations broadly can't break down silos, a support stack stitched together from four or five systems is a harder case, not an easier one. Wrong or outdated answers from a support bot aren't a model failure; they're a data-pipeline failure wearing a chat interface.
What should you ask your AI vendor?
Ask exactly what "trained on your data" means, how quality is measured, and who validated it before it reached your bot.
The report doesn't define what counts as "using data to train AI" — fine-tuning, RAG retrieval, and embeddings are very different exposures, and vendors have incentive to blur them. Push for specifics: which data sources feed responses, how staleness is caught, and what the quality-monitoring process actually looks like, since 62% of organizations in this same survey say they don't have that nailed down internally either.
Frequently asked questions
Does this mean AI support tools are unreliable right now?
Not necessarily — it means reliability depends heavily on the underlying data pipeline, which the survey suggests is unfinished at most organizations in 2026.
Is this specific to customer support software?
No — the Snowflake/Omdia survey covered organizations generally, not support teams specifically, so treat the figures as a directional warning, not a support-sector statistic.
Who ran the survey and who should read it skeptically?
Snowflake, a data-platform vendor, co-produced the research with analyst firm Omdia — worth weighing when interpreting how favorably "using data" is framed.
Source: No Jitter.