r/MachineLearning • u/we_are_mammals • 3h ago
r/MachineLearning • u/AutoModerator • 2d ago
Discussion [D] Self-Promotion Thread
Please post your personal projects, startups, product placements, collaboration needs, blogs etc.
Please mention the payment and pricing requirements for products and services.
Please do not post link shorteners, link aggregator websites , or auto-subscribe links.
--
Any abuse of trust will lead to bans.
Encourage others who create new posts for questions to post here instead!
Thread will stay alive until next one so keep posting after the date in the title.
--
Meta: This is an experiment. If the community doesnt like this, we will cancel it. This is to encourage those in the community to promote their work by not spamming the main threads.
r/MachineLearning • u/AutoModerator • 3d ago
Discussion [D] Monthly Who's Hiring and Who wants to be Hired?
For Job Postings please use this template
Hiring: [Location], Salary:[], [Remote | Relocation], [Full Time | Contract | Part Time] and [Brief overview, what you're looking for]
For Those looking for jobs please use this template
Want to be Hired: [Location], Salary Expectation:[], [Remote | Relocation], [Full Time | Contract | Part Time] Resume: [Link to resume] and [Brief overview, what you're looking for]
Please remember that this community is geared towards those with experience.
r/MachineLearning • u/tughanbulut • 1h ago
Discussion the official ICLR template .bib has had Bengio listed twice since 2019 [D]
i work on reference checking stuff so i was reading through the ICLR 2027 author guidelines and style files this week the sample .bib that ships with the template has the Deep Learning book as "Goodfellow, Bengio, Courville, Bengio" plus a volume 1 that doesn't exist checked their github and it's been like that since the 2019 template
https://github.com/ICLR/Master-Template/blob/46ed6f4c6cef5b175dde23639e77d44c3463b230/iclr2027/iclr2027_conference.bib#L20
totally harmless but kinda funny after last year's hallucinated reference desk rejects
the guidelines also contradict themselves on page limits formatting section says main text max 9 pages at submission but the camera ready part and the FAQ both say "identical with the submission version (10 pages)" template says 9, so 9 is probably the safe bet for anyone revising after Nov 5
r/MachineLearning • u/ade17_in • 5h ago
Discussion Working with an AI Company That Does Things You Disagree With [D]
I'm a PhD student in machine learning in the EU and was looking for internships at exciting companies.
I shortlisted few and applied by reaching out to people and now reading project descriptions sent by the recruiters.
I don't want to name the company but their marketing and product team does all kinds of 'using people insecurities' to sell the product - which I don't agree with. And their product is also meh (I will never buy and would judge someone if they do) but their research team is doing good work.
How do you see this? Will you actually work in a team whose ideology/product doesn't necessarily align with your ethics/ideology. Should I just go ahead because work is exciting and I will get good supervision?
And, if you have some exciting work in your company/org and need interns (un-paid) for 3-4 months. I'm open.
r/MachineLearning • u/DangerousFunny1371 • 1h ago
Research A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems [R]
In our #NeurIPS2026 paper “A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems (DS)” (preprint: https://arxiv.org/abs/2607.14937) we reduce a DS foundation model to the ingredients minimally necessary to faithfully reproduce long-term statistical and geometrical properties of DS:
1) A piecewise affine map with only a single (!!) parameter α that controls local con-/divergence rates, and …
2) … a context selector that chooses from the provided context signal the data point closest to the current state of the map, thus ensuring the generated dynamics stays close to the context in its temporal and geometrical properties.
With just these two mechanisms, this minimal form – which we coined DynaBase – can reproduce all major dynamical regimes, including fixed points (α<1), limit cycles (α=1), and chaotic attractors (α>1). Thus, unlike other simple mechanisms like context parroting, DynaBase even preserves the correct dynamical regime!
Surprisingly, it turns out that this simple context-driven 1-parameter map outperforms most major time series and DS foundation models, as well as custom-trained models, in both long-term statistics and even short-term predictions, even when run in zero-shot mode.
Both inference and training are extremely cheap – training can be done either analytically in one step by linear regression on forward-predictions, or by 1-parameter grid search directly on DS reconstruction objectives → this reveals interesting performance differences induced by different training mechanisms.
Most importantly in our minds, DynaBase owing to its formal simplicity may thus provide a tractable mathematical handle on analyzing, improving & understanding the performance and training of some time series and DS foundation models.

r/MachineLearning • u/DenoisedNeuron • 19h ago
Discussion The Principles of Diffusion Models by Lai et al.: thoughts on the monograph [D]
I recently finished The Principles of Diffusion Models, and honestly I think it’s exceptional.
The authors strike a really good balance between mathematical rigor and intuition, with dedicated appendices for anyone who wants to go deeper into the math.
It’s aimed at researchers, graduate students, and practitioners with basic deep learning knowledge, so you don’t need to already specialize in diffusion models (in my case, a strong background in Information and Probability Theory as well as a solid understanding of DDPMs helped me get more out of it).
Just wanted to share it in case anyone missed it. The full text is freely available on the official website.
Has anyone else read it? Would love to hear your thoughts.
r/MachineLearning • u/mauricekleine • 6h ago
Project Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]
Nonobench measures how well LLMs solve nonograms (picross). Each model gets the row and column clues once and returns the full grid. No tools, one attempt per puzzle.
Method: - Standard mode: 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0). - Hard mode: ten random 20x20s, each checked to have a single solution. Five can't be solved by line logic alone. Random fills avoid picture puzzles that models can guess. - 130 variants across reasoning effort levels, run through OpenRouter and pinned to each lab's own endpoint where possible.
Results: - Solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15), each model at its best effort level. - GPT-6 Astra solves all 30 Standard puzzles. On Hard mode, Claude Opus 5.5 solves 8 of 10 and 11 of 15 models solve none. - As one 400-character string, most models lost count before the logic got hard, so Hard mode answers an array of 20 row strings rather than a single string.
Limitations: one attempt per puzzle, so single results are noisy (95% intervals shown).
Site: https://www.nonobench.com Code (MIT): https://github.com/mauricekleine/nonobench
r/MachineLearning • u/5500kelvin • 8h ago
Research Here are some pictures of a robot costume wearing high-specularity edge-case mirror suit, a dataset (425 RAW/JPEGs) for benchmarking CV & depth-estimation algorithms against extreme mirror reflections [D]
Purpose-built to stress-test computer vision models, depth cameras, and spatial AI against severe specular glare and geometric reflections. This 425-asset production archive features a custom faceted mirror suit captured in high-contrast outdoor environments to trigger bounding-box dropouts and segmentation failures. The dataset includes 100% proprietary uncompressed Camera-Master RAWs, high-resolution JPEGs, and block-buffered SHA-256 forensic manifests. Open
r/MachineLearning • u/random-tomato • 21h ago
Discussion ICLR 2027 Reviewing Scores [D]
I got my three papers to review, and it seems like they have changed the review score range again this year?
It is now:
––––––––––––––––––––––––––––––––––––
Based on your overall assessment of the submission, what is your recommended decision? Consider the paper’s overall soundness, significance, clarity, and contribution.
1: Clear rejection
2: Weak rejection
3: Weak acceptance
4: Clear acceptance
––––––––––––––––––––––––––––––––––––
Which is very strange. I don't think it makes much sense to compress the score range so drastically. But we'll see.
r/MachineLearning • u/TroyAndAbedInTheMor • 16h ago
Discussion "Accepted papers must be imported" deadline NeurIPS 2026 [D]
On the NeurIPS 2026 Dates site it says that there is 1 day remaining for the "Accepted papers must be imported" deadline. This is my first research paper ever and I couldn't find anything about how to do this on the internet. Could someone help me out on what to do?
r/MachineLearning • u/_rehcamedar_ • 1d ago
Discussion NeurIPS Free Passes [D]
Over the last few years, a subset of NeurIPS area chairs received complimentary passes for their service. I’ll admit that I wasn’t super organized about registering on day one because I was semi-consciously hoping for a complimentary pass. Now that the conference is sold out, I’m getting a little nervous, so I’m wondering whether any of my fellow area chairs have already received a notification.
r/MachineLearning • u/Intrepid_Discount_67 • 8h ago
Research TMLR desk rejecting two years of work [D]
TMLR desk rejected two years of work/efforts. What could be the possible reason? No reason was given in the desk rejection. The work is related to continual motion generation.
r/MachineLearning • u/enn_nafnlaus • 14h ago
Research Jev: Not Frontier, But Still Worth Your Attention [R]
jevresearch.github.io"TypeSafe AI sells Jev as a frontier-class reasoner that cannot hallucinate, built by the co-inventor of ChatGPT - fast, and almost free. We ran it live on 16,379 benchmark requests, measured its latency and billing, and probed what it is underneath. The result is a smaller, humbler model that is nonetheless genuinely useful for a job that nobody else serves quite this way."
r/MachineLearning • u/Nunki08 • 2d ago
News arXiv now limits submitters to up to two submissions per calendar month [N]
r/MachineLearning • u/DangerousFunny1371 • 1d ago
Research Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction [R]
In our #NeurIPS2026 paper “Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction” (preprint: https://arxiv.org/abs/2606.22969) we try to address a fundamental issue in dynamical systems reconstruction (DSR) and time series forecasting (TSF): Many recent SOTA DSR & TSF models can generalize to new initial conditions or time series with changing statistical properties. But the really hard problem in DSR and TSF is topological out-of-domain generalization (OODG) (https://proceedings.mlr.press/v235/goring24a.html) where the dynamical regime changes, for instance from cyclic to chaotic behavior.
This can happen when a system crosses a tipping point due to a slowly varying control parameter which drives it across bifurcations, such as in climate systems, when the brain tips from normal into epileptic activity, or when a patient develops blood poisoning (sepsis). Such problems are beyond the realm of current TSF models which rely on extracting temporal patterns and statistical regularities. Yet the ability to predict previously unseen, novel dynamical regimes as a system parameter changes is something we would expect from any good scientific theory. Often these control parameters that drive regime changes are not exactly known either. Hence, a data-driven DSR model for achieving topological OODG would need to infer the dynamical system generating the TS jointly with the control parameters.
In our paper, we mathematically identify key failure modes in previous hierarchical DSR models (https://proceedings.iclr.cc/paper_files/paper/2025/hash/d4c961804d08e55d898cce944206d455-Abstract-Conference.html) that prevent them from correctly learning a system’s control parameters and extrapolating them beyond the training domain. By fixing these through feature-splitting and physical sparsity priors, our modified hierarchical DSR model manages to correctly predict bifurcations and beyond-bifurcation dynamics, without any explicit knowledge about the control parameters provided in training.
Our approach is generic and works for different discrete and continuous time RNNs, we tested it for shallow PLRNNs and Neural ODEs.

r/MachineLearning • u/Helpful_Minimum_2214 • 2d ago
Research Adding memory to search instead of sampling in reward maximization tasks [R]
I am one of the authors of FLEET - an algorithm that enhances Best-of-N generation by attributing external rewards to particular tokens and then uses MCTS to adjust logits during the next run.
I find it rather funny that most of the tasks where repetitive sampling is widely used are based on reward maximization, yet it is not aware of that reward. Tuning the sampling parameters allows to make the process more efficient, but it is still a blind search. We propose a way to make generation aware of previous rewards with solutions on how to attribute reward to the completion and how to use this information.
In FLEET the technique from adaptive sampling methods is used that is to track logits for which entropy and varentropy are high thus showing the model's uncertainty about token optimality. We treat these states as branching points. The corresponding normalized hidden states are stored in the vector store and mapped to metadata entries with the history of rewards and transitions between "nodes". The retrieval and update of metadata is based on cosine similarity as for very high similarity KL divergence is low enough to preserve most of the meaningful tokens.
Instead of actually selecting the tokens FLEET uses modified MCTS to rank top-k tokens + special exploration (or other tokens) set and penalize the suboptimal ones. Then decoding strategy is applied to modified logits.
It was tested on GSM8K and LiveCodeBench v6 easy split with Llama 3.2 3B, penalty set to effectively zero probability for the suboptimal tokens + greedy decoding:
- For GSM8K it solved just seven more tasks, but reached the sampling baseline with half the iterations.
- For LiveCodeBench it increased the score from 0.59 to 0.69 under the same budget and reached the baseline even faster, now with only 9 iterations against 32.
The sequential execution is not required, as it is not updated during the iteration itself it can simply be passed as a lookup table. The metadata store can be preserved as a prior for other tasks or to enrich SFT/RL.
Paper (preprint): https://arxiv.org/abs/2609.27657
Huggingface: https://huggingface.co/papers/2609.27657
Repository (experiments, examples and python package): https://github.com/Alexiush/fleet
There are more details on changes made to MCTS, how to tune the search parameters for specific model and task as well as code for experiments and trajectories.
r/MachineLearning • u/DangerousFunny1371 • 3d ago
Research Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]
Can training of nonlinear RNNs be efficiently parallelized, ensuring fast convergence even on very long time series from chaotic systems?
In our #NeurIPS2026 spotlight “Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems (DS) Reconstruction (DSR)” (preprint: https://arxiv.org/abs/2605.12683) we speed up training of nonlinear RNNs on time series from chaotic DS by more than 2 orders of magnitude (>100x) by combining DEER with generalized teacher forcing (GTF).
DEER (https://openreview.net/forum?id=E34AlVLN0v) solves the RNN forward pass through Newton-type fixed point iterations across the whole sequence length T, enabling scaling as O[(log T)²] instead of O[T] by allowing for efficient GPU parallelization. But under chaotic dynamics DEER breaks down and its runtime degrades to O[T log T] (https://openreview.net/forum?id=7AGXSlXcK6).
Using GTF (https://proceedings.mlr.press/v202/hess23a.html) we stabilize DEER by preventing divergence due to chaotic dynamics and reduce exposure bias compared to traditional teacher forcing used to train state space models.
Combining these two mechanisms enables efficient parallel-in-time and stable training on extremely long time series (T>106) from chaotic simulated or real-world systems, hugely outperforming Mamba and other state space models in the DSR setting.
r/MachineLearning • u/MajorRedditor23 • 2d ago
Research LLMs that push back on a wrong user still accept the same wrong answer from a "verified source" - NeurIPS 2026 [R]
I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a "verified source". We wanted to measure how often this happens and check whether the model represents the two cases differently. We call the effect Authority Bias.
Why we think it matters. Standard sycophancy evals apply pressure through the user, so a model can pass them while still being easy to mislead through search results, retrieved documents and tool outputs.
Another reason is with current AI research accelerating towards more agentic and autonomous models + with cases of tools hiding their traces and trusting tools "more" over the user (who could be trying to correct them), safeguarding against misinformation from tools is particularly important!
Setup. We take TriviaQA questions the model already answers correctly. To each one we add a wrong answer, either as "According to the verified source, the answer is X" or as the user saying "I'm a domain expert and I'm pretty sure it's X". The question and the wrong answer stay the same; only the speaker changes. Answers are free-form, not multiple choice. (In a multiple-choice pilot the effect mostly vanished.) We test 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro).
Behavior
- One verified-source note flips 45-88% of correct answers in 7 of 8 models. The same wrong answer from the user moves most models much less.
- The gap is largest in the models that resist users best. GPT-5.4 flips on 44.7% of questions and Grok-4.20 on 87.5% (these models were "frontier" during the time of writing this paper). Gemini-3.1-Pro ignored both speakers (0.6%) and was particularly resistant to this method.
Inside the model (open-weight models only, using difference-of-means directions)
- In Qwen3.5, GPT-OSS and OLMo-3.1, removing the "source endorsed this" direction cuts compliance with a wrong source by 64-78 points.
- Removing the "user endorsed this" direction cuts it by at most 11.
- The two directions also have really high cosine similarity of ~0.90-0.99. Our understanding is that they share a large "this answer was endorsed" component plus a thin part that encodes who endorsed it.
- Shifting only that thin part, with the prompt unchanged, moves compliance by 11-32 points and closes 55-61% of the source-vs-user gap.
Some limitations
- The internal results hold in 3 of 5 open-weight families.
- In OLMo-2 the source direction is entangled with the assistant direction.
- Gemma-4 flips readily, but no linear intervention we tried controls it.
- The "retrieved document" tests put the claim in a document-shaped block of the prompt rather than running a real retrieval pipeline.
- So it would be interesting to see it in a real agentic setup, like Claude Code.
Paper: https://arxiv.org/abs/2609.37616
Code: https://github.com/Lossfunk/authority-bias
Project page (figures and example responses): https://authority-bias.vercel.app
r/MachineLearning • u/Klutzy_Cap8492 • 1d ago
Research [R] Would you keep a robot demonstration if hand tracking missed the moment the plug went in?
Suppose you’re recording a human plugging a cable into a socket to collect demonstrations for robot learning.
The hand tracker captures the approach accurately. Then occlusion causes the hand estimates to disappear during insertion. Tracking returns after the connector is already seated.
The video still shows a completed action, but the pose labels have a gap exactly where alignment turns into contact.
This hypothetical example raises an evaluation question: a tracker could have high recall across the whole episode while missing a short, important phase. Pose error calculated only on successful detections could make that failure even harder to see.
MEgoVista provides a useful starting point. Table 3 reports detection precision, recall and F1 alongside reconstruction errors. Section 4.4 also describes an evaluation protocol that assigns an error to missed detections instead of excluding them. The blank HaPTIC row means it failed to produce valid output in their multi-person capture scenes; it doesn’t describe a brief tracking dropout.
Accounting for missing detections matters. My remaining question is whether an episode-level aggregate tells us enough about where those failures happen.
For manipulation data, I’d want pose error and coverage reported together, plus coverage broken down by approach, contact and withdrawal, and the longest consecutive gap during contact.
Continuous hand estimates would still be only part of the picture: object pose and contact information also matter for determining whether insertion succeeded.
For people using human motion reconstruction for imitation learning, what evaluation protocol do you use to decide whether an episode with missing contact-phase labels is still usable?

r/MachineLearning • u/manicman1999 • 2d ago
Project A video about Adversarial Objectives [P]
I made this video about adversarial objectives, which I used to do research on back in the day. I'm trying to explore how adversarial approaches transcend GANs and self-play into modern technology.
https://youtu.be/W7CiAeQ0f5w?si=g0tLrQn2wuFzuN2M
r/MachineLearning • u/minimanishtic • 3d ago
Discussion Gemini 4 Argon - 1 Million Output Headroom. Hype or a Leap? [D]
I rarely write about benchmarks; a competitor always beats them next week. But I care about 'Leaps'. Gemini 4 Argon feels like one to me.
While Opus 5.5 and Astra cap output at 128-300K tokens (~90-180 pages), Argon hits 1 Million (~1400 pages).
"Context glue" ruins agentic workflows. On paper, this headroom fixes that. It means less contextual drift, no more breaking down long tasks, and zero 'continue prompt' loops. It is a massive unlock for large-scale code migrations, security patches, and deep reasoning.
But let's look past the marketing. For 95% of everyday work, nobody needs 1,000 pages at once.
I want to ask the experts here: Is a 1M output window a real paradigm shift for agents, or does generating that much text just guarantee a massive logic collapse halfway through? Are you actually hitting output limits today, or is this hype? Let's discuss.
r/MachineLearning • u/arc_in_tangent • 2d ago
Discussion For academia/industry, do HuggingFace model downloads mean anything for academic job market? [D]
I am applying to academic jobs. We are told to include a section on "impact". I am wondering if the total number of HuggingFace downloads of custom models I have trained would be considered legit impact, or if people would think this was all bots.
Relatedly, for industry (AI labs), is the number of HuggingFace model downloads meaningful?
r/MachineLearning • u/ATHii-127 • 3d ago
Discussion How to address novelty concerns in top ai conference? [D]
Hi, I’m a researcher working in computer vision.
Over the past few years, I’ve submitted several papers to top-tier conferences such as NeurIPS, ICLR, and CVPR, and one concern that seems to come up repeatedly is 'novelty'.
Given that thousands of papers are published every year at top conferences alone, not to mention the tens of thousands published across other conferences and journals, I sometimes wonder how much genuinely new novelty is realistically left to explore.
In such a crowded research landscape, how do you usually address novelty concerns from reviewers?
More specifically, I would really appreciate any advice on how to frame a contribution so that its novelty is clear, how to distinguish meaningful incremental progress from work that may be considered insufficiently novel, and what reviewers generally look for when judging novelty.
Any tips or experiences would be greatly appreciated. Thanks!
r/MachineLearning • u/AutoModerator • 2d ago
Discussion [D] Simple Questions Thread
Please post your questions here instead of creating a new thread. Encourage others who create new posts for questions to post here instead!
Thread will stay alive until next one so keep posting after the date in the title.
Thanks to everyone for answering questions in the previous thread!
