AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns
Muhammad Hassan, Ramazan Yener, Ece Gumusel, Masooda Bashir
Read on arXiv →Key claim
User experience failures in AI chatbots impact health information access.
In plain English
Imagine you're trying to get health information quickly and easily through a chatbot. These AI systems are supposed to help, but many users find themselves frustrated. They might struggle to access the service, face unreliable responses, or have a poor experience interacting with the chatbot. Sometimes, they even run into issues with billing or feel their privacy is at risk. This is what's called access barriers and service unreliability, which can lead to negative experiences for users.
In this study, the authors looked at over 15,000 reviews from various AI healthcare chatbots to understand these problems better. They found that users frequently reported issues related to access, usability, and trust. By framing these chatbots as part of a larger information infrastructure, the authors highlight how these failures can significantly impact users' experiences.
What’s new here is the focus on specific breakdowns in user experience and how they relate to the overall effectiveness of these chatbots. This research offers actionable insights for designers and policymakers, suggesting that improving access, interaction quality, and addressing privacy concerns can lead to better digital health systems. For anyone building or improving AI healthcare chatbots, these findings emphasize the importance of user experience and trust.
The study provides new insights into user experiences with AI healthcare chatbots, highlighting specific breakdowns.
The analysis is based on a substantial dataset of user reviews, supporting the claims made.
Deep reliability assessment
The methodology supports a large-scale descriptive map of what dissatisfied app-store reviewers explicitly complain about across AI healthcare chatbot apps. It is weaker evidence for claims about actual user well-being, true population prevalence, causal impact of chatbot failures, or implicit privacy/security concern because the data are self-selected reviews and the privacy analysis is keyword-based.
Reproducibility
No open-source code or released dataset is mentioned in the provided text. The data source is public Google Play and Apple App Store reviews collected in late 2025, but app-store reviews are dynamic and the exact app list, scraping method, preprocessing choices, and topic-modeling parameters would be needed for strong reproducibility.
Key figure
Figure 1 shows the percentage distribution of the three complaint categories across Android and iOS reviews, comparing access barriers/service unreliability, billing/customer support, and user experience/AI interaction quality by platform.
