Introduction
Gemini Deep Research takes a broad question, proposes a research plan, searches and analyzes sources, then produces a report with links readers can inspect. Google introduced it inside Gemini Advanced in December 2024. It is a research workflow in the Gemini app, powered by Gemini models, rather than a separate product with its own founder. Google’s launch announcement described a system that revises its searches as it learns and lets users edit the plan before research begins.
Google has not disclosed how many people use or pay for Deep Research specifically. Its closest published usage figure is for the wider Gemini app: Google said on August 11, 2026, that it had surpassed 1 billion monthly users. Alphabet separately reported 350 million paid subscriptions across its services in its first-quarter 2026 earnings remarks. Neither figure is a Deep Research subscriber count.
Our team’s hands-on evaluation, completed September 25, 2026, placed Gemini Deep Research second among six AI tools for information and research. The question for a prospective user is whether its broad sourcing and Google connections outweigh the time, limits, and checking that a full research report requires.

Performance
Gemini Deep Research scored 8.87/10 overall. Here is how it performed across all seven factors in our scorecard; the weights are those supplied with the evaluation.
| Evaluation Factor | What It Measures | Weight |
|---|---|---|
| AI Intelligence | Reasoning, planning, tools, context, adaptation, and analysis. | 15% |
| Speed | Consistent delivery time for verified research outputs, including iterations and revisions. | 10% |
| Information and Research Quality | Source discovery, selection, freshness, citations, synthesis, gaps. | 25% |
| Prompt Accuracy | Scope adherence, completeness, accuracy, uncertainty, and avoiding unsupported claims. | 20% |
| User Rating | Weighted public satisfaction from credible review platforms | 10% |
| Review Confidence | Editorial calibration of verification, relevance, recency, diversity, and product fit. | 5% |
| Features & Usability | Search, sources, files, citations, workflows, accessibility, and limits. | 15% |
Breadth is the main strength. Gemini’s 9.3 Information and Research Quality score was the highest in our comparison, just above ChatGPT’s 9.2. Its 9.2 Features & Usability score reflects a useful Google-native workflow: Search is a default source; users can select connected Gmail or Drive content, upload files, add Notebook sources, and edit the proposed plan before starting. Gmail and Drive require the Google Workspace connection to be enabled in Gemini. Users can also deselect Search when a task should use only selected sources. Google explains these controls in its workflow guide.
That combination suits a wide-ranging brief that needs both current web information and context from a team’s Google files. For example, a researcher could ask for a market overview, include internal planning documents from Drive, and revise Gemini’s plan to exclude an irrelevant region before it begins. The resulting report gives the reader a starting structure and sources to investigate further.
Time is a real trade-off. Google says a report usually takes about 5–10 minutes to generate. Its research guide notes that more complex reports can take longer. Our 7.5 Speed score measures the broader path to a usable, checkable answer, including any verification and revisions; it is not a timed measurement of that 5–10 minute claim. Deep Research makes sense when a substantial report is worth waiting for. A quick factual lookup is better handled with a faster tool or a standard search.
Prompt fidelity also affects the result. Gemini scored 8.8 in Prompt Accuracy, compared with ChatGPT’s 9.1. Readers should specify the date range, permitted sources, exclusions, required comparisons, and what to do when evidence is missing. Then they should open the sources behind consequential claims: a relevant citation does not always establish the exact wording of a report sentence.
The underlying model and the research workflow are separate considerations. Google’s product overview discusses Gemini 3 in Deep Research, while its plan page advertises access to Gemini 3.1 Pro and Deep Research. Its model guide gives the clearest choice a user can rely on: all users can make reports with Thinking; Google AI Pro and Ultra users can choose Pro. Those public descriptions do not establish that every report, plan, and region uses an identical model version.
Output features also depend on the setup. Google’s help page says Ultra reports may include in-report charts, diagrams, animations, and interactive simulators. Currently, those in-report visuals are unavailable when a Google Workspace service such as Gmail or Drive is included as a research source. Report-request and simultaneous-request limits vary by plan, and Google says access may change with demand.
Independent research adds perspective but doesn’t determine our score. The DeepResearch Bench evaluated 100 research tasks using measures of report quality and citation behavior. An earlier Gemini 2.5 Pro Deep Research version led its overall research measure at 48.88, ahead of OpenAI Deep Research at 46.98. It averaged 111.21 effective citations, but Perplexity Deep Research had higher citation accuracy: 90.24% versus Gemini’s 81.44%. That distinction matters to a reader who values the precision of each citation more than the number of sources gathered. The benchmark tested earlier systems, so its ordering should not be projected onto every current version.
A June 2026 consulting study tested agents on 70 expert-written tasks with checks for data integrity, reasoning, and deliverable quality. Its acceptance rates were low: 15.7% for OpenAI o3 deep-research and 12.9% each for Gemini 3.1 Pro deep-research and Claude Opus 4.6. The pairwise acceptance differences were statistically indistinguishable. As a preprint focused on difficult consulting work, it warns against relying on an unattended report for a decision, rather than offering a universal product ranking.
Comparisons Before You Buy
Our September 25 scorecard puts the products and their best uses in this order. Factor columns use the same 10-point scale as the performance table.
| Product | Overall Score | AI Intelligence | Speed | Information and Research Quality | Prompt Accuracy | User Rating | Review Confidence | Features & Usability | Best For |
|---|---|---|---|---|---|---|---|---|---|
| ChatGPT | 8.94 | 9.3 | 7.1 | 9.2 | 9.1 | 9.4 | 7.5 | 9.3 | Complex sourced reports from files |
| Gemini Deep Research | 8.87 | 9.1 | 7.5 | 9.3 | 8.8 | 9.1 | 7.5 | 9.2 | Google Workspace integrated research briefs |
| Claude | 8.73 | 9.4 | 7.7 | 8.7 | 9.1 | 9.3 | 7.5 | 8.3 | Nuanced synthesis of long documents |
| Perplexity | 8.72 | 8.5 | 9.2 | 8.8 | 8.6 | 9.2 | 7.5 | 8.7 | Fast current cited information lookups |
| Gemini Notebook | 8.45 | 8.1 | 8.5 | 8.6 | 8.5 | 9.5 | 6.2 | 8.5 | Study selected documents and notes |
| Scite AI | 7.84 | 7.3 | 8.2 | 7.8 | 7.8 | 9.1 | 6.0 | 7.7 | Citation context and literature auditing |
Gemini’s slight research-quality lead over ChatGPT did not give it the highest overall score. ChatGPT led in intelligence, prompt accuracy, user rating, and features; Gemini led in research quality and speed. Both received 7.5 for Review Confidence. That shared score reflects a limitation in the outside reviews: ratings of the general apps do not isolate their named research modes. It restrains confidence in either product’s app-rating evidence, but does not explain the gap between their two scores.
Claude has the stronger case when nuanced reading of long documents matters most: it scored 9.4 for AI Intelligence, the highest in the group. Perplexity’s 9.2 Speed makes it attractive for current, cited lookups. Gemini Notebook, formerly NotebookLM, fits a selected collection of notes and documents; Scite AI fits work centered on how research papers cite one another. Different research priorities could reasonably produce a different winner.
Why did we use Google Play’s 46.9 million figure for Gemini? Our September 25 public-rating scorecard records the following listings:
| Product | Source used | Public Rating | Votes & Reviews |
|---|---|---|---|
| ChatGPT | Google Play Store | 4.5/5 | 61.8M reviews |
| Gemini Deep Research | Google Play Store | 4.4/5 | 46.9M reviews |
| Claude | Google Play Store | 4.4/5 | 753K+ reviews |
| Perplexity | G2 | 4.4/5 | 368+ reviews |
| Gemini Notebook | Google Play Store | 4.6/5 | 325K+ reviews |
| Scite AI | Trustpilot | 4.4/5 | 274+ reviews |
For the Gemini row, 46.9 million was the Google Gemini app’s displayed headline review count, as recorded in our materials. We selected Google Play as a broad user-satisfaction signal because Deep Research lives in that app, the listing is directly relevant to the interface a user opens, and its sample is exceptionally large. The visible five-star scale and count also make the observation straightforward to inspect. It remains an app-level proxy, not a survey of 46.9 million Deep Research users and not a test of citation quality.
The additional September 25 ratings audit supplied for this update records a discrepancy worth preserving. It lists 4.6/5 and 45.3 million in the Gemini listing’s ratings section, alongside 46.9 million in the page headline. Our scorecard image records 4.4/5 and 46.9 million. The supplied records do not establish why their star ratings differ, so we retain the original scorecard figures as its dated snapshot rather than treating them as the live rating. On September 26, the listing displayed 4.6/5 and 47 million headline reviews.
We also consider other review populations. The supplied September 25 audit records Google Gemini at 4.4/5 from 604 reviews on G2 ratings, 8.5/10 from 171 reviews and ratings on TrustRadius ratings, and 4.5/5 from 77 ratings on Gartner ratings, where 92 reviews were displayed separately. Those business-oriented sources offer useful written context and more relevant professional users for some purchasing decisions. Google Play’s much larger sample is better read as broad app satisfaction. Other sites can capture different experiences, including support problems; their scores are not inherently wrong because they differ.
Public ratings did not determine our hands-on ranking. The scorecard’s User Rating factor uses two platforms under its stated rule, and Review Confidence accounts for sample and product match. Neither factor turns a parent-app rating into evidence that Deep Research itself produces reliable citations.
For a U.S. personal account, these were the relevant plan terms checked September 26, 2026. Google’s limits guide describes compute-based limits that vary with use and can change; the “×” access comparisons are plan-wide descriptions, not guaranteed numbers of Deep Research reports.
Google announced a separate Ultra tier at about $200/month with higher general Gemini limits. Its current plan page states that Ultra plans are billed monthly only. Google published the U.S. Plus price in its Plus announcement and the two Ultra tiers in its subscription update.
Google AI Pro is our recommendation for a typical weekly researcher who will use its higher limits and Google connections. Its verified annual price offers real savings over monthly billing: $19.99 × 12 = $239.88, compared with $199.99 annually, a difference of $39.89. The annual payment works out to about $16.67 per month. Check the terms at purchase because taxes, availability, and offers can affect the amount paid. Google’s plan details show both Pro billing options.
Start with Free for an occasional report. Plus is a modest step up if free access becomes restrictive but Pro’s capacity is unnecessary. Choose Pro for regular, substantial briefs, especially those involving many files or Workspace context. Consider Ultra only when report volume, limits, or its eligible visual output justify a much higher monthly price, and remember that its in-report visuals currently cannot be combined with Gmail or Drive sources.

Conclusion
Gemini Deep Research is particularly good at a wide-ranging brief that also needs Google Workspace context. Search, optional connected sources, uploaded files, Notebook sources, and an editable plan make its research workflow useful; our 9.3 Research Quality score reflects that breadth. ChatGPT’s 8.94 overall score remained ahead of Gemini’s 8.87, and readers who prioritize long-document nuance or fast lookups have credible alternatives in Claude and Perplexity.
Its main practical cost is time and checking. Google documents a usual 5–10 minute wait for a report, with complex work taking longer. Model choice, connected sources, visual output, and usage limits depend on the plan and setup. For regular research, Google AI Pro is the balanced paid option. For occasional work, try the free access first and inspect the claims and citations before relying on the report.
FAQ
Why did Gemini Deep Research rank second overall despite receiving the highest research-quality score?
Gemini scored 9.3 for Information and Research Quality, slightly above ChatGPT’s 9.2. However, ChatGPT scored higher in AI Intelligence, Prompt Accuracy, User Rating, and Features & Usability. As a result, ChatGPT finished at 8.94 overall, while Gemini Deep Research scored 8.87. Gemini is especially strong for broad, source-rich research, but the overall ranking also rewards reasoning, instruction-following, usability, and user evidence.
Can Gemini Deep Research combine web research with private Google files?
Yes. Users can combine Search with uploaded files, Notebook sources, and connected Gmail or Drive content. Search is enabled by default, but users can turn it off when the task should use only selected sources. Gmail and Drive require Google Workspace connections to be enabled. Users can also edit Gemini’s proposed research plan before the report begins, which helps remove irrelevant regions, sources, or subquestions.
Does the 4.4/5 rating from 46.9 million reviews measure Deep Research itself?
No. The figure comes from the broader Google Gemini app, where Deep Research is offered. It is an app-level signal covering the wider Gemini experience, including access, speed, reliability, interface, and subscription friction. It is not a survey of Deep Research users and does not measure citation accuracy. The rating and count also changed across the supplied records, so treat 4.4/5 from 46.9 million reviews as the dated September 25 scorecard snapshot.
How long does a Gemini Deep Research report take, and what limitations should users expect?
Google describes a usual report-generation time of 5 to 10 minutes, with complex reports taking longer. The article’s 7.5 Speed score covers the full path to a usable answer, including verification and revisions, rather than measuring only Google’s generation estimate. All users can create reports with Thinking, while Google AI Pro and Ultra users can choose Pro. Ultra reports may include charts, diagrams, animations, and interactive simulators, but those visuals are currently unavailable when Gmail or Drive is used as a source. Limits also vary by plan, demand, and setup.
Which Google AI plan is the best choice for Deep Research?
Google AI Pro is the best balance for a typical weekly researcher. It costs $19.99 per month or $199.99 per year, includes expanded Deep Research access, the Pro report option, a 1M-token context window, and 5TB of storage. The annual plan saves $39.89 compared with twelve monthly payments and works out to about $16.67 per month. Free is suitable for occasional reports, Plus is a modest upgrade when free limits become restrictive, and Ultra, starting at $99.99 per month, is mainly for users whose report volume, higher limits, or eligible visual features justify the additional cost.