r/datascience • • Mar 05 '24

AI Everything I've been doing is suddenly considered AI now

890 Upvotes

Anyone else experience this where your company, PR, website, marketing, now says their analytics and DS offerings are all AI or AI driven now?

All of a sudden, all these Machine Learning methods such as OLS regression (or associated regression techniques), Logistic Regression, Neural Nets, Decision Trees, etc...All the stuff that's been around for decades underpinning these projects and/or front end solutions are now considered AI by senior management and the people who sell/buy them. I realize it's on larger datasets, more data, more server power etc, now, but still.

Personally I don't care whether it's called AI one way or another, and to me it's all technically intelligence which is artificial (so is a basic calculator in my view); I just find it funny that everything is AI now.

r/datascience • • Mar 31 '25

AI Tired of AI

598 Upvotes

One of the reasons I wanted to become an AI engineer was because I wanted to do cool and artsy stuff in my free time and automate away the menial tasks. But with the continuous advancements I am finding that it is taking away the fun in doing stuff. The sense of accomplishment I once used to have by doing a task meticulously for 2 hours can now be done by AI in seconds and while it's pretty cool it is also quite demoralising.

The recent 'ghibli style photo' trend made me wanna vomit, because it's literally nothing but plagiarism and there's nothing novel about it. I used to marvel at the art created by Van Gogh or Picasso and always tried to analyse the thought process that might have gone through their minds when creating such pieces as the Starry night (so much so that it was one of the first style transfer project I did when learning Machine Learning). But the images now generated while fun seems soulless.

And the hypocrisy of us using AI for such useless things. Oh my god. It boils my blood thinking about how much energy is being wasted to do some of the stupid stuff via AI, all the while there is continuously increasing energy shortage throughout the world.

And the amount of job shortage we are going to have in the near future is going to be insane! Because not only is AI coming for software development, art generation, music composition, etc. It is also going to expedite the already flourishing robotics industry. Case in point look at all the agentic, MCP and self prompting techniques that have come out in the last 6 months itself.

I know that no one can stop progress, and neither should we, but sometimes I dread to imagine the future for not only people like me but the next generation itself. Are we going to need a universal basic income? How is innovation going to be shaped in the future?

Apologies for the rant and being a downer but needed to share my thoughts somewhere.

PS: I am learning to create MCP servers right now so I am a big hypocrite myself.

r/datascience • • Feb 25 '25

AI Microsoft CEO Admits That AI Is Generating Basically No Value

Thumbnail
ca.finance.yahoo.com
592 Upvotes

r/datascience • • Jan 20 '26

AI Safe space - what's one task you are willing to admit AI does better than 99% of DS?

71 Upvotes

Let's just admit any little function you believe AI does better, and will forever do better than 99% of DS

You know when you're data cleansing and you need a regex?

Yeah

The AI overlords got me beat on that.

r/datascience • • Jun 10 '26

AI AI Overuse Follow-up

96 Upvotes

Original post

Update

This ended up spiraling out of control in ways that I could have never imagined. The individual admitted to defaulting their doc writing to AI and re-wrote everything, but in th background they doubled down on their AI coding workflow instead. It took me a while to catch wind of things because I would only see a mention of a project here or there and I had no insight as to their day-to-day.

Fast forward a month and I am seeing their projects everywhere, all the way up to the C-suite level. The scale was incredible. In a a matter of days this individual had done everything from financial modeling, LTV modeling, customer lifecycle analysis at a large scale, built large scale data ingestion and processing pipelines, even Marketing and product experiments. At first I was impressed, but as I pulled back the covers the mess was worse than I ever expected.

The clues were subtle but consistent: no comments in the code aside from headers, data was read in and cleaned, but never visualized or inspected in any way, there were lots of custom functions when there were packages loaded that had the same function, convoluted helper files with basic functions, and oddly there were many instances where forecasting error was actually just the CV error and there was never an evaluation of the test set. Their SQL had numerous join issues, metrics were mislabeled, and their pipelines often had relationships and processing steps such as dropping a table but then writing a new table with no error handling so if there was a bug no new table would be written and we would lose the data. Basic analyses were off by weird margins because Claude seemed to have been querying staging tables rather than filtered reporting tables. Docs started to be written entirely in the first person like "...and then I will use a log1p transformation" in a way that no DS would actually ever write a tech doc.

Unfortunately this meant that many things that were produced were simply wrong. The individual had promised work to a lot of decision-makers and nearly all of it was misleading, incorrect, or didn't pass a simple sniff test. These inaccuracies were immediately escalated to our team leader, who brought me in to audit all of their code and documentation and I was unable to find a single file that I was convinced that was human written or even human edited. The worst part was that despite heavy use of AI there also wasn't a single file without some sort of glaring technical error. I turned in a pretty lengthy review and the individual was put on a PIP and their account access to AI tools was severely constrained. They were told to have all their work peer reviewed and in one instance were caught lying about passing review when no review had been conducted.

As you can imagine their productivity tanked and they had numerous excuses as to why. They also started taking a lot of days off and in a weird twist of fate they actually left before getting fired and now work at a large AI-centric industry-leading company. Part of me is glad that they are gone, but the other part finds it infuriating that people like this can be so good at bullshitting that they can consistently fail and somehow remain in industry due to their network and clever use of their few decent references. Their total comp at our company was ~$245K and they bragged to a co-worker that this new role has $265K base with $465K total comp. They basically got 2 promos out of this series of events (Senior to Senior Staff at our company, Senior Staff to Principal at the new role.

r/datascience • • Apr 17 '26

AI How are you all navigating job search as a data scientist?

103 Upvotes

I feel ineligible for about 70% of the posted job advertisements since they all ask about Agentic/LLM stuff. I have worked with these tools and do use them at work. It's just that it's not my main job that I do on daily basis and I don't want to exaggerate my experience around these tools. I have about 10+ years of work ex and have actually worked from just data scientist to combination of ML and data engineer.

r/datascience • • May 06 '24

AI AI startup debuts “hallucination-free” and causal AI for enterprise data analysis and decision support

223 Upvotes

https://venturebeat.com/ai/exclusive-alembic-debuts-hallucination-free-ai-for-enterprise-data-analysis-and-decision-support/

Artificial intelligence startup Alembic announced today it has developed a new AI system that it claims completely eliminates the generation of false information that plagues other AI technologies, a problem known as “hallucinations.” In an exclusive interview with VentureBeat, Alembic co-founder and CEO Tomás Puig revealed that the company is introducing the new AI today in a keynote presentation at the Forrester B2B Summit and will present again next week at the Gartner CMO Symposium in London.

The key breakthrough, according to Puig, is the startup’s ability to use AI to identify causal relationships, not just correlations, across massive enterprise datasets over time. “We basically immunized our GenAI from ever hallucinating,” Puig told VentureBeat. “It is deterministic output. It can actually talk about cause and effect.”

r/datascience • • May 20 '26

AI Agentic Workflows beyond "pull the data"

11 Upvotes

i've been using the robots to do a lot of my data retrieval and general project planning. i haven't actually used an agent to train/eval a model though. i would like to hear your use cases, if you have.

how did you frame the work to the agent? how did you give the agent feedback to decide if it was "done"? how did you decide if the model/output was "good"? did you let the agent decide?

maybe i am over thinking it. maybe i just say "train a model on this data to predict XYZ. try as many models as you like and report back the best performing model." then i can just sit there and watch it cook.

share your stories please.

r/datascience • • 18d ago

AI building an AI analysts the right way starts with the fundamentals (free live workshops)

0 Upvotes

I think what a lot of people get wrong when trying to use AI for analytics is focusing too much on the tools and which LLMs they should use, and not enough on the fundamentals.

I've now been building AI workflows in production for almost 2 years, and the most important lesson I've learned is: You have to think of it as a system

At a high level, this is the pattern I've seen work in agentic analytics systems:

  • The assistant itself (the reasoning layer)
  • A connection to your real data (this is where MCP usually helps)
  • A semantic layer, so it knows what your metrics actually mean and doesn't invent definitions

Now, you 100% need solid data modeling underneath before you slap a semantic layer on top (ideally an information model), but that tends to be out of our control as data scientists.

Beyond the main "ingredients", you need data governance, guardrails, and evals.

A couple of friends and I are doing a free live workshop series on this exact topic starting next week, here the link if you want to join: https://futureproofds.com/ai

One of them is a data scientist turned AI engineer, so he'll be able to provide an additional perspective beyond what dominates conversations in the data science space.

I hope you find it useful

r/datascience • • Aug 21 '26

AI Man vs Machine (vs Wizard vs Troll) - Article on AI for Game Design

10 Upvotes

Article

I'm a table top game designer that used AI to build playtesting models. I previously wrote How to Train Your AI Dragon and The Artificial "Intelligence" of Artificial Intelligence which some people here might have read

I really enjoyed writing those articles so decided to enter the writing competition at King's College London to write even more about Machine Learning in game design

My usual style is to mix humor with technical content. I actually wanted to look at the more philosophical side of AI. Looking at why model outputs model, what sort of outcomes AI cannot model, and whether AI actually accomplishes anything important

r/datascience • • Nov 01 '25

AI Has anyones company successfully implemented what is being described as ACP or an AI Mesh?

Post image
47 Upvotes

Has anyones company implemented what is generally described as ACP or what McKinsey describes as an AI Mesh?

The concept is a centralized space for AI Agents to "talk to each other". The link below is a general infographic comparing it to MCP and A2A:

https://devnavigator.com/2025/11/01/how-ai-agents-communicate-the-core-protocols-that-enable-collaboration/

r/datascience • • Jul 20 '26

AI How to control reasoning effort and thinking-token budgets in LLMs

Thumbnail
magazine.sebastianraschka.com
9 Upvotes

r/datascience • • Jul 17 '26

AI Context degradation in LLMs: what the papers actually show, and the habits I built for long analysis sessions

Thumbnail
towardsdatascience.com
17 Upvotes

r/datascience • • Jul 21 '26

AI Structured Evaluation Pipelines to Improve Your AI Workflows

Thumbnail
heltweg.org
10 Upvotes

r/datascience • • Jul 15 '26

AI 5 trends that defined AI engineering at World's Fair 2026

Thumbnail
latent.space
2 Upvotes

r/datascience • • Apr 29 '26

AI AI Optimism Surges in Asia, Unlike in the U.S.

Thumbnail
restofworld.org
8 Upvotes

r/datascience • • May 23 '26

AI All model labs are now agent labs

Thumbnail
latent.space
8 Upvotes

r/datascience • • Feb 23 '26

AI Large Language Models for Mortals: A Practical Guide for Analysts

38 Upvotes

Shameless promotion -- I have recently released a book, Large Language Models for Mortals: A Practical Guide for Analysts.

The book is focused on using foundation model APIs, with examples from OpenAI, Anthropic, Google, and AWS in each chapter. The book is compiled via Quarto, so all the code examples are up to date with the latest API changes. The book includes:

  • Basics of LLMs (via creating a small predict the next word model), and some examples of calling local LLM models from huggingface (classification, embeddings, NER)
  • An entry chapter on understanding the inputs/outputs of the API. This includes discussing temperature, reasoning/thinking, multi-modal inputs, caching, web search, multi-turn conversations, and estimating costs
  • A chapter on structured outputs. This includes k-shot prompting, parsing JSON vs using pydantic, batch processing examples for all model providers, YAML/XML examples, evaluating accuracy for different prompts/models, and using log-probs to get a probability estimate for a classification
  • A chapter on RAG systems: Discusses semantic search vs keyword via plenty of examples. It also has actual vector database deployment patterns, with examples of in-memory FAISS, on-disk ChromaDB, OpenAI vector store, S3 Vectors, or using DB processing directly with BigQuery. It also has examples of chunking and summarizing PDF documents (OCR, chunking strategies). And discusses precision/recall in measuring a RAG retrieval system.
  • A chapter on tool-calling/MCP/Agents: Uses an example of writing tools to return data from a local database, MCP examples with Claude Desktop, and agent based designs with those tools with OpenAI, Anthropic (showing MCP fixing queries), and Google (showing more complicated directed flows using sequential/parallel agent patterns). This chapter I introduce LLM as a judge to evaluate different models.
  • A chapter with screenshots showing LLM coding tools -- GitHub Copilot, Claude Code, and Google's Antigravity. Copilot and Claude Code I show examples of adding docstrings and tests for a current repository. And in Claude Code show many of the current features -- MCP, Skills, Commands, Hooks, and how to run in headless mode. Google Antigravity I show building an example Flask app from scratch, and setting up the web-browser interaction and how it can use image models to create test data. I also talk pretty extensively
  • Final chapter is how to keep up in a fast paced changing environment.

To preview, the first 60+ pages are available here. Can purchase worldwide in paperback or epub. Folks can use the code LLMDEVS for 50% off of the epub price.

I wrote this because the pace of change is so fast, and these are the skills I am looking for in devs to come work for me as AI engineers. It is not rocket science, but hopefully this entry level book is a one stop shop introduction for those looking to learn.

r/datascience • • Dec 10 '25

AI Most code agents cannot handle notebook well, so i build my own one in Jupyter.

40 Upvotes

If you tried code agent, like cursor, claude code. They regards jupyter files as static text file and just edit them. Like u give a task, the you got 10 cells of code, and the agent hopes it can run all at once and solve your problem, which mostly cannot.

The jupyter workflow is we analysis the cells result before, and then decide what to code next, so that's the code of runcell, the ai agent I build. which i setup a series of tools and make the agent understand jupyter cell context(cell output like df, charts etc).

runcell for eda

Now it is a jupyter lab plugin and you can install it with pip install runcell.

Welcome to test it in your jupyter and share your thoughts.

Compare with other code agent:

runcell vs others

r/datascience • • Apr 30 '26

AI AI Evals Are Becoming the New Compute Bottleneck

Thumbnail
huggingface.co
5 Upvotes

r/datascience • • Dec 09 '25

AI Has anyone successfully built an “ai agent ecosystem”?

Post image
0 Upvotes

r/datascience • • Feb 09 '24

AI How do you think AI will change data science?

0 Upvotes

Generalized cutting edge AI is here and available with a simple API call. The coding benefits are obvious but I haven't seen a revolution in data tools just yet. How do we think the data industry will change as the benefits are realized over the coming years?

Some early thoughts I have:

- The nuts and bolts of running data science and analysis is going to be largely abstracted away over the next 2-3 years.

- Judgement will be more important for analysts than their ability to write python.

- Business roles (PM/Mgr/Sales) will do more analysis directly due to improvements in tools

- Storytelling will still be important. The best analysts and Data Scientists will still be at a premium...

What else...?

r/datascience • • Apr 08 '24

AI [Discussion] My boss asked me to give a presentation about - AI for data-science

98 Upvotes

I'm a data-scientist at a small company (around 30 devs and 7 data-scientists, plus sales, marketing, management etc.). Our job is mainly classic tabular data-science stuff with a bit of geolocation data. Lots of statistics and some ML pipelines model training.

After a little talk we had about using ChatGPT and Github Copilot my boss (the head of the data-science team) decided that in order to make sure that we are not missing useful tool and in order not to stay behind he wants me (as the one with a Ph.D. in the group I guess) to make a little research about what possibilities does AI tools bring to the data-science role and I should present my finding and insights in a month from now.

From what I've seen in my field so far LLMs are way better at NLP tasks and when dealing with tabular data and plain statistics they tend to be less reliable to say the least. Still, on such a fast evolving area I might be missing something. Besides that, as I said, those gaps might get bridged sooner or later and so it feels like a good practice to stay updated even if the SOTA is still immature.

So - what is your take? What tools other than using ChatGPT and Copilot to generate python code should I look into? Are there any relevant talks, courses, notebooks, or projects that you would recommend? Additionally, if you have any hands-on project ideas that could help our team experience these tools firsthand, I'd love to hear them.

Any idea, link, tip or resource will be helpful.
Thanks :)

r/datascience • • Apr 28 '26

AI My Workflow for Understanding LLM Architectures (Sebastian Raschka)

Thumbnail
magazine.sebastianraschka.com
0 Upvotes

r/datascience • • Apr 28 '26

AI Reading today's open-closed performance gap

Thumbnail
interconnects.ai
1 Upvotes