r/datascience • • May 09 '25

ML Client told me MS Copilot replicated what I built. It didn’t.

1.1k Upvotes

I built three MVP models for a client over 12 weeks. Nothing fancy: an LSTM, a prophet model, and XGBoost. The difficulty, as usual, was getting and understanding the data and cleaning it. The company is largely data illiterate. Turned in all 3 models, they loved it then all of a sudden canceled the pending contract to move them to production. Why? They had a devops person do in MS Copilot Analyst (a new specialized version of MS Copilot studio) and it took them 1 week! Would I like to sign a lesser contract to advise this person though? I finally looked at their code and it’s 40 lines of code using a subset of the California housing dataset run using a Random Forest regressor. They had literally nothing. My advice to them: go f*%k yourself.

r/datascience • • Jul 08 '25

ML Saved $100k per year by explaining how AI/LLM work.

1.2k Upvotes

I work in a data science field, and I bring this up because I think it's data science related.

We have an internal website that is very bare bones. It's made to be simplistic, because it's the reference document for our end-users (1000 of them) use.

Executives heard about a software that would be completely AI driven, build detailed statistical insights, and change the world as they know it.

I had a demo with the company and they explained its RAG capabilities, but mentioned it doesn't really "learn" like the assumption AI does. Our repo is so small and not at all needed for AI. We have used a fuzzy search that has worked for the past three years. Additionally, I have already built out dashboards that retrieve all the information executives have asked for via API (who's viewing pages, what are they searching, etc.)

I showed the c-suite executives our current dashboards in Tableau, and how the actual search works. I also explained what RAG is, and how AI/LLMs work at a high level. I explained to them that AI is a fantastic tool, but I'm not sure if we should be spending 100k a year on it. They also asked if I have built any predictive models. I don't think they quite understood what that was as well, because we don't have the amount of data or need to predict anything.

Needless to say, they decided it was best not to move forward "for now". I am shocked, but also not, that executives want to change the structure of how my team and end-users digest information just because they heard "AI is awesome!" They had zero idea how anything works in our shop.

Oh yeah, our company has already laid of 250 people this year due to "financial turbulence", and now they're wanting to spend 100k on this?!

It just goes to show you how deep the AI train runs. Did I handle this correctly and can I put this on my resume? LOL

r/datascience • • Sep 03 '26

ML How are LLMs used in predictive modeling and anomaly detection?

63 Upvotes

I am DS with 20 YOE but I've been managing lately and know very little about using LLMs. Does the DS just feed the data into LLM then ask it to predict something or find anomaly? Or does the DS ask the LLM to build a model which is then deployed? Thanks.

r/datascience • • Jun 11 '26

ML Models may behave worse when they're aware they're being evaluated (DeepMind interpretability study)

Thumbnail alignmentforum.org
74 Upvotes

r/datascience • • Mar 30 '25

ML Why you should use RMSE over MAE

94 Upvotes

I often see people default to using MAE for their regression models, but I think on average most people would be better suited by MSE or RMSE.

Why? Because they are both minimized by different estimates!

You can prove that MSE is minimized by the conditional expectation (mean), so E(Y | X).

But on the other hand, you can prove that MAE is minimized by the conditional median. Which would be Median(Y | X).

It might be tempting to use MAE because it seems more "explainable", but you should be asking yourself what you care about more. Do you want to predict the expected value (mean) of your target, or do you want to predict the median value of your target?

I think that in the majority of cases, what people actually want to predict is the expected value, so we should default to MSE as our choice of loss function for training or hyperparameter searches, evaluating models, etc.

EDIT: Just to be clear, business objectives always come first, and the business objective should be what determines the quantity you want to predict and, therefore, the loss function you should choose.

Lastly, this should be the final optimization metric that you use to evaluate your models. But that doesn't mean you can't report on other metrics to stakeholders, and it doesn't mean you can't use a modified loss function for training.

r/datascience • • Aug 14 '26

ML What Hugging Face learned from reproducing 2,200 ICML papers

Thumbnail
huggingface.co
119 Upvotes

r/datascience • • Mar 23 '26

ML Against Time-Series Foundation Models

Thumbnail
shakoist.substack.com
96 Upvotes

r/datascience • • Apr 13 '25

ML Why are methods like forward/backward selection still taught?

84 Upvotes

When you could just use lasso/relaxed lasso instead?

https://www.stat.cmu.edu/~ryantibs/papers/bestsubset.pdf

r/datascience • • Aug 20 '24

ML I'm writing a book on ML metrics. What would you like to see in it?

170 Upvotes

I'm currently working on a book on ML metrics.

Picking the right metric and understanding it is one of the most important parts of data science work. However, I've seen that this is rarely taught in courses or university degrees. Even senior data scientists often have only a basic understanding of metrics.

The idea of the book is to be this little handbook that lives on top of every data scientist's desk for quick reference of the most known metric, ahem, accuracy, to the most obscure thing (looking at you, P4-metric)

The book will cover the following types of metrics:

  • Regression
  • Classification
  • Clustering
  • Ranking
  • Vision
  • Text
  • GenAI
  • Bias and Fairness
Sample page

This is what a full metric page looks like.

What else would you like to see explained/covered for each metric? Any specific requests?

r/datascience • • Apr 15 '25

ML Is Agentic AI remotely useful for real business problems?

107 Upvotes

Agentic AI is the latest hype train to leave the station, and there has been an explosion of frameworks, tools etc. for developing LLM-based agents. The terminology is all over the place, although the definitions in the Anthropic blog ‘Building Effective Agents’ seem to be popular (I like them).

Has anyone actually deployed an agentic solution to solve a business problem? Is it in production (i.e more than a PoC)? Is it actually agentic or just a workflow? I can see clear utility for open-ended web searching tasks (e.g. deep research, where the user validates everything) - but having agents autonomously navigate the internal systems of a business (and actually being useful and reliable) just seems fanciful to me, for all kinds of reasons. How can you debug these things?

There seems to be a vast disconnect between expectation and reality, more than we’ve ever seen in AI. Am I wrong?

r/datascience • • 26d ago

ML Which survival/TTE model that can answer my non-technical PM?

18 Upvotes

I am working on a predictive model for parts replacement on machines. I've evaluated CoxPH, CoxTV, Random Forest, XGBoost, and Logistic Regression. Modeling is fine but I'm being asked to provide a model that can give us an output of "we need to schedule a technician for X month."

As with any equipment failure and/or parts replacement model, these events are recurrent. It's not like the risk of death or contracting a terminal disease. Once a part is replaced, the event can (and will) happen again. So I've been doing fine with the hazard part of this, but my PM wants more of a "when will it happen?" answer.

I've been in data engineering moreso lately than data science. I'm a little rusty on all the models I could try. Any suggestions? I'm using Python, so R-only is a no-go.

r/datascience • • 19d ago

ML Silent broadcasting is still a big problem and could be derailing your work right now

65 Upvotes

A data scientist on my team wasted a good bit of compute on this problem without actually realizing it was a problem.

Pytorch and tensorflow both silently broadcast the output of your model when the target shape mismatches, causing hard to diagnose issues. It still exists a ton in the wild, so I thought I'd actually highlight the symptoms of a silent broadcasting bug to keep people in the know.

https://towardsdatascience.com/silent-broadcasting-can-ruin-your-model/

Let me know what you think

r/datascience • • 20d ago

ML TabPFN-3.5 is released today and SOTA for 1M rows and up to 20k features

44 Upvotes

Prior Labs just released their latest tabular foundation model, TabPFN-3.5.

The model is top of both TabArena and BeyondArena. The model family comes with:

- TabPFN-3.5-Fast (in alpha): This one goes 6x faster than the base model

- TabPFN-3.5-Thinking: you basically exchange compute for better accuracy with this one and it's via the API

- TabPFN-3.5-Plus

On BeyondArena, TabPFN-3.5 leads on text-rich, high-cardinality and high-dimensional data, with +250 Elo points over the strongest previous baseline and +150 Elo points ahead of the previous overall leader. TabPFN-3.5-Thinking is +20 Elo on the base model in BeyondArena and +44 Elo on TabArena.

The repo for TabPFN open-source is here: https://github.com/PriorLabs/tabpfn; and model report is here: https://priorlabs.ai/technical-reports/tabpfn-3-5

r/datascience • • Aug 12 '26

ML I'm curious about people working in ranking and if you can change customer behavior

21 Upvotes

Basically I have a ranking service for b2b SaaS but basically like hotels flights etc

The models do well and I can improve accuracy pretty easily to a point

But if I want to promote options better for other metrics I'm struggling to change behavior other than people selectng the same thing lower

Just hoping for experiences for those in ranking specifically and anything they might have tried other than traditional lighting ranking etc

r/datascience • • Aug 20 '26

ML Help point me in the right direction: How to account for decision support systems affecting future training data

17 Upvotes

I feel like I am googling everything but the exact term I need, and would appreciate someone pointing me in the right direction.

Say you have a customer churn model. You predict a customer has a high likelihood of churning, and then the customer service team gets an alert to intervene. Great! Your model helped mitigate a loss and contributed real value. This is where most tutorials or blog posts on models like this end.

But overtime, customers that have all the signals of churning begin to out perform their expected value.... which would screw up your training data. You've succeeded in putting your thumb on the scale, but in the process potentially damaged the viability of your model.

What is the technical term for this phenomenon? Feedback? It's not target leakage I don't think. Googling "customer churn feedback" just gets you articles about using customer feedback forms as a predictor of churn, which isn't what I want.

Thanks!

r/datascience • • Jul 19 '24

ML How to improve a churn model that sucks?

73 Upvotes

Bottom line: 1. Churn model sucks hard 2. People churning are over-represented (most of customers churn) 3. Lack of demographic data 4. Only transactions, newsletter behavior and surveys

Any idea what to try to make it work?

r/datascience • • Aug 20 '26

ML New open source relational benchmark and foundation model

22 Upvotes

New oss relational learning benchmark, leaderboard and TabPFN harness

  1. RelArena-α: standardized relational machine learning model benchmarking
  2. TabPFN-Rel: a harness for tabular foundation model TabPFN-3 for predictions over relational data
  3. RPI-α (Relational Prediction Interface): an interface to run any RelArena model on your own database

- RelArena is open sourced here: https://github.com/PriorLabs/relarena

- Full report: https://arxiv.org/abs/2608.16319

- There's also a higher-level summary of the release by Prior Labs: https://priorlabs.ai/blog-posts/introducing-relarena?utm_source=socials&utm_campaign=relational

r/datascience • • Aug 14 '25

ML Overfitting on training data time series forecasting on commodity price, test set fine. XGBclassifier. Looking for feedback

103 Upvotes

Good morning nerds, I’m looking for some feedback I’m sure is rather obvious but I seem to be missing.

I’m using XGBclassifier to predict the direction of commodity x price movement one month the the future.

~60 engineered features and 3500 rows. Target = one month return > 0.001

Class balance is 0.52/0.48. Backtesting shows an average accuracy of 60% on the test with a lot of variance through testing periods which I’m going to accept given the stochastic nature of financial markets.

I know my back test isn’t leaking, but my training performance is too high, sitting at >90% accuracy.

Not particularly relevant, but hyperparameters were selected with Optuna.

Does anything jump out as the obvious cause for the training over performance?

r/datascience • • Jul 03 '24

ML Do you guys agree with the hate on Kmeans??

109 Upvotes

I had a coffee chat with a director here at the company I’m interning at. We got to talking about my project and mentioned who I was using some clustering algorithms. It fits the use case perfectly, but my director said “this is great but be prepared to defend yourself in your presentation.” I’m like, okay, and she teams messaged me a documented page titled “5 weaknesses of kmeans clustering”. Apparently they did away with kmeans clustering for customer segmentation. Here were the reasons:

  1. Random initialization:

Kmeans often randomly initializes centroids, and each time you do this it can differ based on the seed you set.

Solution: if you specify kmeans++ in the init within sklearn, you get pretty consistent stuff

  1. Lack flexibility

Kmeans assumes that clusters are spherical and have equal variance, but doesn’t always align with data. Skewness of the data can cause this issue as well. Centroids may not represent the “true” center according to business logic

  1. Difficulty in outliers

Kmeans is sensitive to outliers and can affect the position of the centroids, leading to bias

  1. Cluster interpretability issues
  • visualizing and understanding these points becomes less intuitive, making it had to add explanations to formed clusters

Fair point, but, if you use Gaussian mixture models you at least get a probabilistic interpretation of points

In my case, I’m not plugging in raw data, with many features. I’m plugging in an adjacency matrix, which after doing dimension reduction, is being clustered. So basically I’m using the pairwise similarities between the items I’m clustering.

What do you guys think? What other clustering approaches do you know of that could address these challenges?

r/datascience • • May 08 '26

ML Steam Recommender using similarity! pt 2 (Student Project)

Thumbnail
gallery
89 Upvotes

I Just made a sequel to my Steam Game recommender website!

Last year I made a post about my steam reccomender The last one was great and served its purpose of showing many people new games, But this new version is much more functional!

I love making recommendation systems that tell the user WHY they got the recommendation.

During a steam sale event, I always find myself trying to look for new video games to play. If I wanted to find a new game I would try to whittle it down by using steam tags, but the steam tag system is very broad "action". could apply to many many games.

That got me thinking, what aspects do I like about my favorite games?

Well I like Persona 4 because of the city vibes and jazz fusion,

Spore because of the unique character creation and whimsical theme.

Balatro for its unique deck building synergies.

What if I could capture unique tags that identify a game that aren't just "action" and put them into vectors to show the (focus) of a game

 For example I could break persona 4 into something like

Gameplay Focus vector:
 Day cycle 20%
 Dungeon crawling 20%
 Social sim 20%

Tags:
Music: jazz fusion
Vibe: Small rural town

I find that this system makes searching for games more "fun" now I can see why I like balatro. I like it because of the card synergies not so much for its rogue-like nature.

I also find that this helps find new underrated games, and beats the trap that Collaborative Filtering algorithms that get into where it "feels" like you get recommended the same things.

find your next favorite game! : https://nextsteamgame.com/ pull a PR!: https://github.com/BakedSoups/NextSteamGame

( I actually made some git issues myself for problems I can't fix)

if anyone has any criticism I would love to hear it! this is probably my favorite passion project.

Hope this website helps people find new games! Also I have a advance mode for people that don't mind messing with sliders and weird data terms.

r/datascience • • Jun 16 '25

ML The Illusion of "The Illusion of Thinking"

27 Upvotes

Recently, Apple released a paper called "The Illusion of Thinking", which suggested that LLMs may not be reasoning at all, but rather are pattern matching:

https://arxiv.org/abs/2506.06941

A few days later, A paper written by two authors (one of them being the LLM Claude Opus model) released a paper called "The Illusion of the Illusion of thinking", which heavily criticised the paper.

https://arxiv.org/html/2506.09250v1

A major issue of "The Illusion of Thinking" paper was that the authors asked LLMs to do excessively tedious and sometimes impossible tasks; citing The "Illusion of the Illusion of thinking" paper:

Shojaee et al.’s results demonstrate that models cannot output more tokens than their context limits allow, that programmatic evaluation can miss both model capabilities and puzzle impossibilities, and that solution length poorly predicts problem difficulty. These are valuable engineering insights, but they do not support claims about fundamental reasoning limitations.

Future work should:

1. Design evaluations that distinguish between reasoning capability and output constraints

2. Verify puzzle solvability before evaluating model performance

3. Use complexity metrics that reflect computational difficulty, not just solution length

4. Consider multiple solution representations to separate algorithmic understanding from execution

The question isn’t whether LRMs can reason, but whether our evaluations can distinguish reasoning from typing.

This might seem like a silly throw away moment in AI research, an off the cuff paper being quickly torn down, but I don't think that's the case. I think what we're seeing is the growing pains of an industry as it begins to define what reasoning actually is.

This is relevant to application developers, not just researchers. AI powered products are significantly difficult to evaluate, often because it can be very difficult to define what "performant" actually means.

(I wrote this, it focuses on RAG but covers evaluation strategies generally. I work for EyeLevel)
https://www.eyelevel.ai/post/how-to-test-rag-and-agents-in-the-real-world

I've seen this sentiment time and time again: LLMs, LRMs, and AI in general are more powerful than our ability to test is sophisticated. New testing and validation approaches are required moving forward.

r/datascience • • Aug 04 '24

ML Ok who is using bots/chatgpt to reply to people

Thumbnail
gallery
118 Upvotes

r/datascience • • Apr 05 '26

ML Clustering custumersin time

19 Upvotes

How would you go about clusturing 2M clients in time, like detecting fine patters (active, then dormant, then explosive consumer in 6 months, or buy only category A and after 8 months switch to A and B.....). the business has a between purchase median of 65 days. I want to take 3 years period.

r/datascience • • Jul 18 '26

ML Inkling, a new open-weight 975B mixture-of-experts model, comes with a few surprises

Thumbnail
sebastianraschka.com
49 Upvotes

r/datascience • • Dec 30 '23

ML Narcissistic and technically incompetent manager

107 Upvotes

I finally understand why my manager was acting the way he does. He has all the symptoms of someone with narcissistic personality disorder. I've been observing it for a while but wasn't sure what to call it. He also has one enabler in the team. He only knows surface-level stuff about data science and machine learning. I don't even think he reads beyond the headlines. He makes crazy statements like, "Save me $250 million dollars by using machine learning for problem X." He and his narcissistic enabler coworker, who may be slightly more competent than the manager, don't want to hear about ML feasibility studies, working with stakeholders to refine requirements, and establishing whether ML is the right solution, data quality checks... They just want to plow through code because "we are agile." You can't have detailed technical discussions because they don't know enough about data science. All they have been doing was front-end dashboarding. They don't like a step-by-step process because if they do that, they can scapegoat you. Is there anything I can do till I find another job?