WildSeekEvaluating Language Models for Information-Seeking
WildSeek is a manually annotated dataset of 3,077 information-seeking queries taken from real user–LLM conversations. Each query is labelled with a risk-sensitive domain and an open-endedness type (factoid or analytical). We used it to train classifiers and ran them over 1.8M+ real queries. Alongside the dataset, we propose an evaluation framework for LLMs in the context of information seeking, which scores responses for reliability, fairness and safety. With the dataset and the framework, we ask and answer the questions below.
How often do people use LLMs to seek information?
~40%
of user turns are information-seeking.
Details
Information seeking is the most common use: it accounts for about 40% of user turns, from 42% in LMSYS-Chat-1M to 74% in SES.
How much of that information seeking is risk-sensitive?
37%
of information-seeking queries touch a risk-sensitive domain.
Details
Of the information-seeking queries, 37% touch a risk-sensitive domain, led by moral values & religion, health, and economic & financial. 60% are analytical rather than factoid.
Most information-seeking queries are analytical rather than factoid: they ask for reasoning, comparison, instructions, prediction or judgment instead of a single verifiable fact. Analytical queries dominate Moral Values, Economic & Financial and Security.
Yes. Analytical queries fail more often than factoid ones.
Details
Failures rise in 4 of 6 criteria, with significant results for at least one model in each. The gap is largest for overreliance: 21.3% vs. 7.3%, about 3× higher, for every model with and without search. Anthropomorphism shows the same pattern (6.2% vs. 1.9%).
FactoidAnalytical
Failure rate (%), averaged over models and setups (Table 5). * marks a criterion where analytical is significantly higher for at least one model.
What goes wrong most often?
1Overreliance14.3%
2Vulnerable population9.3%
3Sycophancy8.6%
Average failure rate across models and setups.
Details
The most frequent failures are overreliance, vulnerable-population harm, US bias and sycophancy. Overreliance averages 14.3% across models and setups. The other three average about 9% each.
Web search significantly improves factuality only for Gemini-3.1 (+0.049). It does not reduce US-centric framing overall, and only Gemini-3.1 becomes consistently safer with it. Claude Sonnet 4.6 becomes more sycophantic with search (from 10.3% to 48.5% on analytical queries).
The model citing the most credible sources draws on the fewest web domains.
Details
With search on, GPT-5.4 cites the largest share of highly credible sources (about 76–77%). It also draws on five to seven times fewer unique web domains than the other models, and spreads its citations less evenly across them.
These are 3,077 information-seeking queries sampled from four corpora of real human–LLM conversations. Annotators labelled each one by hand with the risk-sensitive domain it touches and whether it asks for a fact or for analysis.
Domains where a wrong or incomplete answer could affect someone's life, safety, health, money or decisions.
7 classes3 annotatorsFleiss κ 0.82
Domain definitions
Layer 3
Open-endedness
Factoid queries can be answered from a single verifiable source. Analytical queries need reasoning, comparison, instructions, prediction or judgment.
2 classes3 annotatorsFleiss κ 0.62majority vote
Decision rules
Factoid
Definitions, historical or scientific facts, descriptions of entities or systems, text lookup. Typical signals: What is / Who is / When did.
Analytical
Procedural steps, comparisons and trade-offs, predictions, advice or value judgments. Typical signals: How to / Why / Should I / Best / Compare.
Borderline
If one verifiable fact or a concise description answers it, it is factoid. Instructions, predictions and opinions are always analytical.
What users ask, in the wild
Our ModernBERT classifiers labelled every turn in four in-the-wild corpora (1.8M+ prompts). The figures below reproduce Figures 2–4 of the paper. See performance of classifiers
User intent by dataset
Share of each dataset's queries per intent category. Darker cells are larger shares.
Risk-sensitive domains among information-seeking queries
Share of information-seeking queries per domain (N = classified queries).
Factoid vs. analytical, by domain
Analytical queries dominate Moral Values, Economic & Financial and Security.