That's the third product to use "Deep Research" in its name.
The first was Gemini Deep Research: https://blog.google/products/gemini/google-gemini-deep-resea... - December 11th 2024
Then ChatGPT Deep Research: https://openai.com/index/introducing-deep-research/ - February 2nd 2025
Now Perplexity Deep Research: https://www.perplexity.ai/hub/blog/introducing-perplexity-de... - February 14th 2025.
Just a side note: The Wikipedia page for "Deep Research" only mentions OpenAI – https://en.wikipedia.org/wiki/Deep_Research
This is bizarre, wasn't Google the one who claimed the name and did it first?
Gemini was also "use us through this weird interface and also you can't if you're in the EU"; that + being far behind OpenAI and Anthropic for the past year means, they failed to reach notoriety, partly because of their own choices.
Honestly I don‘t get why everybody is saying Gemini is far behind. Like for me Gemini Flash Thinking Experimental performs far far better then o3 mini
Seconding this. I get really great results from Flash 2.0 and even Pro 1.5 for some things compared to OpenAI models.
And their 2.0 Thinking model is great for other things. When my task matters, I default to Gemini.
I find the problem with Gemini is the rate limits. Really constrictive.
There's a lot of mental inertia combined with an extremely fast moving market. Google was behind in the AI race in 2023 and a good chunk of 2024. But they largely caught up with Gemini 1.5, especially the 002 release version. Now with Gemini 2 they are every bit as much of a frontier model player as OpenAI and Anthropic, and even ahead of them in a few areas. 2025 will be an interesting year for AI.
Arguably Google is ahead. They have many non-llm uses (waymo/deepmind etc) and they have their own hardware, so not as reliant on Nvidia.
Demis Hassabis isn't very promotional. The other guys make more noise.
It varies a lot for me. One day it takes scattered documents, pasted in, and produces a flawless summary I can use to organize it all. The next, it barely manages a paragraph for detailed input. It does seem like Google is quick to respond to feedback. I never seem to run into the same problem twice.
> It does seem like Google is quick to respond to feedback.
I'm puzzled as to how that would work, when people talk about quick changes in model behavior. What exactly is being adjusted? The model has already been trained. I would think it's just randomness.
Magic
And fine tuning.
Choose your fighter...
High level overview: https://www.datacamp.com/tutorial/fine-tuning-large-language...
More detail: https://www.turing.com/resources/finetuning-large-language-m...
Nice charts: https://blogs.oracle.com/ai-and-datascience/post/finetuning-...
The big platforms also seem to employ an intermediate step where they rewrite your prompt. I've downloaded my ChatGPT data and found substantial changes from what I wrote. Usually for the better. Changes to the way it rewrites changes the results.
System prompts have a huge impact on output. Prompts for ChatGPT/etc are around a thousand words, with examples of what to do and what not to do. Minor adjustments there can make a big difference.
I've found this as well. On a good day Gemini is superb. But otherwise, awful. Really weird.
It was far behind. That's what I kept hearing on the Internet until maybe a couple weeks ago, and it didn't seem like a controversial view. Not that I cared much - I couldn't access it anyway because I am in the EU, which is my main point here: it seems that they've improved recently, but at that point, hardly anyone here paid it any attention.
Now, as we can finally access it, Google has a chance to get back into the race.
o3 mini is still behind o1 pro, it didn't impress me.
I think the people who think anybody is close to OpenAI don't have pro subscription
The $200 version? It's interesting that it exists, but for normal users it may as well... not. I mean, pro is effectively not a consumer product and I'd just exclude it from comparison of available models until you can pay for a single query.
o3-mini isn't meant to compete with o1, or o1 pro mode.
It’s speed makes it better for me to iterate … o1 pro is just too slow or not yet good enough to wait 5 minutes…
I can tell you why I just stopped using Gemini yesterday.
I was interested in getting simple summary data on the outcome of the recent US election and asked for an approximate breakdown of voting choices as a function age brackets of voters.
Gemini adamantly refused to provide these data. I asked the question four different ways. You would think voting outcomes were right up there with Tiananmen Square.
ChatGPT and Claude were happy to give me approximate breakdowns.
What I found interesting is that the patterns if voting by age are not all that different from Nixon-Humphrey-Wallace in 1968.
Gemini's guardrails are unnecessarily strict. As you mentioned, there's a topical restriction on election-related content, and another where it outright refuses to process images containing anything resembling a face. I initially thought Copilot was bad in this regard—it also censors election-related questions to some extent, but not as aggressively as Gemini. However, Gemini's defensiveness on certain topics is almost comical. That said, I still find it to be quite a capable model overall.
I think somebody has read your comment and fixed it...
Elicit AI just rolled out a similar feature, too, specifically for analyzing scientific research papers:
https://support.elicit.com/en/articles/4168449
I find it better for my phd topic actually. Its paper recommendations are quite well.
It is a term of art now in the field.
Is there a problem with this if it's not trademarked? It's like saying Apple Maps is the nth product called "Maps".
I, for one, am glad they are standardising on naming of equivalent products and wish they would do it more (eg. "reasoning" vs "thinking", "advanced voice mode" vs "live")
Not a trademark lawyer, but I don’t think Deep Research qualifies for trademark protection because it is “merely descriptive” of the product’s features. The only way to get a trademark like that is through “acquired distinctiveness”, but that takes 5 years of exclusive use and all these competitors will make that route impossible.
https://www.emergentmind.com also offers Deep Research on ArXiv papers (experimental)
I own DeepCQ.com since early 2023 - Which could do "deepseek" for financial research. Maybe I just throw this on the pile, too.
It failed my first test which concerned Upside magazine. All of these deep research versions have failed to immediately surface the most famous and controversial article from that magazine, "The Pussification of Silicon Valley." When hinted, Perplexity did a fantastic job of correcting itself, the others struggled terribly. I shouldn't have to hint though, as that requires domain knowledge that the asker of a query might be lacking.
We're mere months into these things, though. These are all version 1.0. The sheer speed of progress is absolutely wild. Has there ever been a comparable increase in the ability of another technology on the scale of what we're seeing with LLMs?
I wouldn’t go so far as to say it was definitely faster, but the development of mobile phones post-iPhone went pretty quick as well.
> pussification of silicon valley upside magazine
Google nor bing can find this
https://www.google.com/search?q=pussification+of+silicon+val...
Nothing with "pussification" in the title for me there.
Wild. My results are literally dozens of posts about the article.
https://imgur.com/a/1hTJVkl
I don't see the article you are mentioning
Wild. My results are literally dozens of posts about the article.
https://imgur.com/a/1hTJVkl
About the article, not any link to the article itself.
It is possible that the original article is no longer accessible online.
The only link I have found is a reproduction of the article[1], but I am unable to access the full text due to a paywall. I no longer have access to academic resources or library memberships that would provide access.
My Google search query was:
which returned exactly one result.
I suspect the article's low visibility in standard Google searches, requiring operators like 'inurl:', might be because its PageRank is low due to insufficient backlinks.
[1] https://www.proquest.com/docview/217963807?sourcetype=Trade%...
Can't find it either.
I see a reference to the comment, a guiardian article about the article but not the article itself.
Perhaps it’s softnuked in the eu or something?
Do you have Google SafeSearch or Bing's equivalent turned on perhaps?
I reckon it might be triggered by the word 'pussification' to refuse to return any results related to that.
If you're using a corporate account, it's possible that your account manager has enabled SafeSearch, which you may not be able to disable.
Local censorship laws, such as those in South Korea, might also filter certain results.
My standard prompts when I want thoroughness:
"Did you miss anything?"
"Can you fact check this?"
"Does this accurately reflect the range of opinions on the subject?"
Taking the output to another LLM with the same questions can wring out more details.
I'd expect a "deep research" product to do this for me.
You forgot Huggingface researchers - https://www.msn.com/en-us/news/technology/hugging-face-resea...
and BTW - I post an exact same spirit comment an hour ago... So I guess Today's copycat ethics aren't solely for products- but also for comment section . LOL.
Your comment from earlier wasn’t as easy to digest as this one. I don’t think that person copied you at all.
Thanks. I accept the criticism of being less digest and more opinionated. But at the end of the day it provide the same information.
Don't get me wrong - I don't mind to be copied on the Internet :), but I find this behavior quite rude, so I just mentioned it.
Thinking simonw is stealing your comment is comedy moment of the day
Said comment, so other's don't have to dig around in your history:
"Since google, everyone trying replicate this feature... (OpenAI, HF..) It's powerfull yes, so as asking an A.I and let him sythezise all what he fed.
I guess the air is out of the ballon from the big players, since they lack of novel innovation in their latest products."
I'd say the important differences are that simonw's comment establishes a clear chronology, gives links, and is focused on providing information rather than opinion to the reader.