WildSeekEvaluating Language Models for Information-Seeking

WildSeek is a manually annotated dataset of 3,077 information-seeking queries taken from real user–LLM conversations. Each query is labelled with a risk-sensitive domain and an open-endedness type (factoid or analytical). We used it to train classifiers and ran them over 1.8M+ real queries. Alongside the dataset, we propose an evaluation framework for LLMs in the context of information seeking, which scores responses for reliability, fairness and safety. With the dataset and the framework, we ask and answer the questions below.

How often do people use LLMs to seek information?

~40%

of user turns are information-seeking.

Details

Information seeking is the most common use: it accounts for about 40% of user turns, from 42% in LMSYS-Chat-1M to 74% in SES.

See the in-the-wild figures →

How much of that information seeking is risk-sensitive?

37%

of information-seeking queries touch a risk-sensitive domain.

Details

Of the information-seeking queries, 37% touch a risk-sensitive domain, led by moral values & religion, health, and economic & financial. 60% are analytical rather than factoid.

See the domain breakdown →

Do people ask for facts or for analysis?

60%

of information-seeking queries are analytical.

Details

Most information-seeking queries are analytical rather than factoid: they ask for reasoning, comparison, instructions, prediction or judgment instead of a single verifiable fact. Analytical queries dominate Moral Values, Economic & Financial and Security.

See factoid vs. analytical by domain →

Do analytical queries get less safe answers?

+4.0 pp

Yes. Analytical queries fail more often than factoid ones.

Details

Failures rise in 4 of 6 criteria, with significant results for at least one model in each. The gap is largest for overreliance: 21.3% vs. 7.3%, about 3× higher, for every model with and without search. Anthropomorphism shows the same pattern (6.2% vs. 1.9%).

Factoid Analytical

Failure rate (%), averaged over models and setups (Table 5). * marks a criterion where analytical is significantly higher for at least one model.

What goes wrong most often?

  1. 1Overreliance14.3%
  2. 2Vulnerable population9.3%
  3. 3Sycophancy8.6%

Average failure rate across models and setups.

Details

The most frequent failures are overreliance, vulnerable-population harm, US bias and sycophancy. Overreliance averages 14.3% across models and setups. The other three average about 9% each.

Explore failure rates →

Does web search make answers better?

Not reliably

Only Gemini-3.1 clearly benefits from search.

Details

Web search significantly improves factuality only for Gemini-3.1 (+0.049). It does not reduce US-centric framing overall, and only Gemini-3.1 becomes consistently safer with it. Claude Sonnet 4.6 becomes more sycophantic with search (from 10.3% to 48.5% on analytical queries).

See reliability results →

Are credible sources also diverse?

Not necessarily

The model citing the most credible sources draws on the fewest web domains.

Details

With search on, GPT-5.4 cites the largest share of highly credible sources (about 76–77%). It also draws on five to seven times fewer unique web domains than the other models, and spreads its citations less evenly across them.

See source diversity →

The WildSeek dataset

These are 3,077 information-seeking queries sampled from four corpora of real human–LLM conversations. Annotators labelled each one by hand with the risk-sensitive domain it touches and whether it asks for a fact or for analysis.

Queries per risk-sensitive domain

queries in total

Explore the annotated queries →
Layer 1

User intent

Is the turn information-seeking at all? There are five classes: information seeking, content creation, coding, no request, not English.

3,907 turns2 annotatorsCohen's κ 0.80

Browse intent labels →
Layer 2

Risk-sensitive domain

Domains where a wrong or incomplete answer could affect someone's life, safety, health, money or decisions.

7 classes3 annotatorsFleiss κ 0.82

Domain definitions
Layer 3

Open-endedness

Factoid queries can be answered from a single verifiable source. Analytical queries need reasoning, comparison, instructions, prediction or judgment.

2 classes3 annotatorsFleiss κ 0.62majority vote

Decision rules
Factoid
Definitions, historical or scientific facts, descriptions of entities or systems, text lookup. Typical signals: What is / Who is / When did.
Analytical
Procedural steps, comparisons and trade-offs, predictions, advice or value judgments. Typical signals: How to / Why / Should I / Best / Compare.
Borderline
If one verifiable fact or a concise description answers it, it is factoid. Instructions, predictions and opinions are always analytical.

What users ask, in the wild

Our ModernBERT classifiers labelled every turn in four in-the-wild corpora (1.8M+ prompts). The figures below reproduce Figures 2–4 of the paper. See performance of classifiers

User intent by dataset

Share of each dataset's queries per intent category. Darker cells are larger shares.

Risk-sensitive domains among information-seeking queries

Share of information-seeking queries per domain (N = classified queries).

Factoid vs. analytical, by domain

Analytical queries dominate Moral Values, Economic & Financial and Security.