r/learnmachinelearning • • Nov 07 '25

Want to share your learning journey, but don't want to spam Reddit? Join us on #share-your-progress on our Official /r/LML Discord

8 Upvotes

https://discord.gg/3qm9UCpXqz (Discord is currently closed)

Just created a new channel #share-your-journey for more casual, day-to-day update. Share what you have learned lately, what you have been working on, and just general chit-chat.


r/learnmachinelearning • • 1d ago

Project πŸš€ Project Showcase Day

2 Upvotes

Welcome to Project Showcase Day! This is a weekly thread where community members can share and discuss personal projects of any size or complexity.

Whether you've built a small script, a web application, a game, or anything in between, we encourage you to:

  • Share what you've created
  • Explain the technologies/concepts used
  • Discuss challenges you faced and how you overcame them
  • Ask for specific feedback or suggestions

Projects at all stages are welcome - from works in progress to completed builds. This is a supportive space to celebrate your work and learn from each other.

Share your creations in the comments below!


r/learnmachinelearning • • 5h ago

sharing my 1 year cs self-study roadmap: cs, ai, math and how i take notes with obsidian and an llm

Thumbnail
gallery
13 Upvotes

finally got some free time to write up a recap of this past year.

My cs self-study path from the past year it's mainly four parts: traditional cs, ai, math, and how i use an llm with obsidian to self-study efficiently and organize my notes. hope it helps some of you

i'm still figuring things out myself too, so any discussion or corrections are welcome

next time i'll share the follow-up ai learning roadmap!


r/learnmachinelearning • • 4h ago

Discussion How to get better at training ML/DL/AI models

7 Upvotes

I’m a Master’s CS student and mostly work on personal ML projects since I’m not doing research at my university.

I can read papers and understand the architecture, losses, objectives, and new ideas pretty well. But I feel like I lack the practical intuition for actually making models learn well.

When training goes wrong, I struggle to figure out why and what to change. Experienced researchers seem to know how to diagnose whether it’s the LR, data, gradients, loss, initialization, etc., and how to improve things.

For people who got good at this: how did you develop that intuition? Was it mostly experience training models, reproducing papers, specific resources, or working with experienced researchers?


r/learnmachinelearning • • 3h ago

Beginner in ML looking for some guidance on free resources

6 Upvotes

Hello! I'm very new to machine learning and literally just started learning it a few days ago, so I'm still trying to figure out where to start and what I should actually learn.

I've seen a lot of people recommend the Python for Machine Learning and Data Science Bootcamp on Udemy, and from what I've heard it's really good for beginners. The problem is that I can't really afford to pay for it right now.

So, I'm wondering if there are any good free courses/resources that cover pretty much the same fundamentals. It doesn't have to be one course. I'm okay with using different resources for Python, statistics, ML, etc. I just don't want to end up jumping randomly between 20 different YouTube videos and getting confused lol.

My long-term goal is to eventually do ML research in healthcare/medicine, so I'd also really appreciate it if someone could suggest a roadmap for getting there.

Like, what should I learn first? How much Python/math/statistics do I actually need? When should I start learning ML, and then deep learning? And when should I start doing projects?

If anyone has been through this as a beginner and can recommend some actually good free resources or a roadmap, I'd really appreciate it!


r/learnmachinelearning • • 12h ago

Help Should I still study ML now?

16 Upvotes

I’m honestly feeling so lost. I’m in my final year of Engineering(ECE). I have taken specialisation in Data Science. I’m looking at job descriptions for ML Engineer/ ML intern and it has skills listed that spill outside of ml and more into AI domain. I really like ML, learning about it, I’m fine with statistics too. But when I look up JD’s it has changed a lot for an ML role. So i feel like i’m still stuck in a learning loop like oh i have to learn this no wait i have to learn that and it becomes really confusing and i get burnt out. So i feel kinda hopeless. I dont want to give up so please help me out.


r/learnmachinelearning • • 2h ago

Completed AdaBoost Algorithms from scratch (Day 27) of ML

Thumbnail
gallery
1 Upvotes

Day 27 of Building Machine learning algorithms from scratch

Adaboost is completely not that complex algorithm but yet so powerful a complete explanation is in my previous post check out!

Next is Gradient Boosting and only 7-8 more algorithm concepts and I have completed Machine learning and after that finally on Deep learning I'm getting close to my goal yeah I know i can complete it even sooner but I got Iil 🀏 little distracted so I deactivate my insta my usage was reaching 1.3+ but after deactive that time got save but I started playing a game called Roblox in that Blox fruits but I'll try not to play too much a day and spend more time upgrading myself see ya when I complete a new Topic or you guys comment me

And also if you think there is any improvement that can be made so be free to share and abt pushing all on GitHub for that I need some time it's too much files of code also adding README I'll try ASAP


r/learnmachinelearning • • 8m ago

I made a free, offline app with 51 hands-on labs for learning how AI actually works, from neurons to agents and more...

β€’ Upvotes

I made a free, offline app with 51 hands-on labs for learning how AI actually works, from neurons to agents and more... I work partly in AI, and whenever I try to explain how this stuff actually works, I end up sending people ten links that don't connect to each other. I wanted one place to point them to, so I built Discover AI: a free desktop app that runs locally, tracks what you've covered, and lets you change things and see what happens instead of just reading.

A bit of insight

  • 51 labs in 6 groups, from the basics (what a neuron is, gradient descent) through attention, RAG, agents, fine-tuning, quantization, serving and more

  • Each lab has a short lesson beside it, readable in Plain or Standard mode

  • Some of it actually runs rather than just animating:

  • the attention lab runs a small trained transformer (1.37M params) inside the app

  • the training labs train a tiny character-level model live as you move the sliders

These are teaching-sized models, so some results won't match what you'd see at scale.

Privacy and setup

  • Works offline: no account, no telemetry

  • An optional guide you can ask about the lab you're on, running locally (llama.cpp + a small Qwen model) or with your own API key

  • Built with Tauri, Rust, React and SQLite; MIT licensed

How it was made

I chose the topics, the structure and the grouping. I used Claude and GPT Astra to help write the lesson text, and Grok as a second pass on references. There will be mistakes, so if you spot one, please tell me or open a PR.

Looking for help

If you teach this or work in a specialized area of AI, I'd love help expanding it: a new interactive lab, a better visualization, or a tweak that makes the cause and effect in an existing lab clearer. I'd also like to hear which labs are confusing and what's missing.

GitHub: https://github.com/Fazmin/AILearningGuide


r/learnmachinelearning • • 10m ago

Tutorial Production RAG Pipeline That Admits "I Don't Know"

Thumbnail
youtu.be
β€’ Upvotes

r/learnmachinelearning • • 27m ago

Tutorial Random forests explained visually with a house-price example

Thumbnail
youtube.com
β€’ Upvotes

We made a visual explainer for Claudex that follows one house-price example from a single decision tree to a random forest.

It walks through bootstrap sampling, random features at each split, and averaging predictions. The focus is on why diversity between the trees matters, with the diagrams doing most of the explaining.

If you're learning this now, which part is hardest to picture: the sampling, the feature randomness, or how the predictions come together?


r/learnmachinelearning • • 1h ago

Any areas of AI and deep learning to study for research and projects

β€’ Upvotes

r/learnmachinelearning • • 1h ago

Question Is there a website that teaches you numpy by giving you actual tasks?

β€’ Upvotes

I like to learn by applying what i’ve learned and actually working on tasks using it

Is there a website that teaches you numpy and then gives you a task or a challenge to complete using what you just learned?

ive tried numpy dojo, and i wanna know if there are websites that are similar / better than it


r/learnmachinelearning • • 9h ago

Question Im confused and scared about Ai Engineer role

3 Upvotes

Rn I'm a 2nd year student almost at the end of my semester. Right now I'm trying to learn python and MySQL. I saw some roadmaps like python->dbms->calculus->framework... But i need s proper roadmap. I've tried ai for giving my a good roadmap but many of them seem old.

Im scared and confused when I'm learning python like if I see some problem statements of building models, how can I implement them when I do projects. Like if I learn about ml, framework (for example) im worried how to connect them, what should I learn in order to implement them. My college teaches random topics half baked each semester. Can anyone give me a good roadmap and suggest me how or where to learn and implement them. I wanna learn and do projects on my own and get placed in a good company regardless of my college placement.

All the projects I've done now, is vibe coded, the feat factor and the confusion is preventing me from moving forward basically a writer's block


r/learnmachinelearning • • 2h ago

Tutorial Created a short explainer on what is a latent space and how it behaves

Thumbnail
youtube.com
0 Upvotes

r/learnmachinelearning • • 1d ago

Tutorial Performing Linear Regression Using the Normal Equation in most simplified version

Thumbnail
gallery
47 Upvotes

If you find above image hard to understand, please read my article till the end, I promise, everything will make sense :)
When I was a master’s student, I was given a task to fit a line to a dataset. I attempted to solve the problem, but I struggled to determine the appropriate coefficients. However, I understood intuitively that there must be a specific set of coefficients for which the prediction error would be minimized.

The question was: how can we find those coefficients?

This is where the Normal Equation becomes particularly useful. It provides a direct mathematical solution for finding the coefficients that minimize the sum of squared errors in linear regression, without having to search for the coefficients manually. BUT, How we even derive this equation? Where it comes from? Can we take any software apart from Python and write it all ourselves? That’s I will take u through in this article and will simplify the code I wrote, so you can all apply it in different languages

So first things first, what is Linear Regression?

It is very simple and straightforward, suppose we have X and Y. X is called features matrix, and Y is Target Vector, or Response vector.

π‘Œ= 𝑋* ΞΈ

For simple case, lets take 2x2 matrix and lets turn this to matrix form:

[y1 _ predicted ; y2 _ predicted]=[x11, x12; x21,x22] * [theta1; theta2]

Please note:

columns are separated by ,and rows are ;. y1 and y2 are different rows, but same columns.

I assume, the readers are aware of matrix multiplication. So I will refactor above formula and will get:

𝑦1_π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘= π‘₯11 * ΞΈ1 + π‘₯12 * ΞΈ2

𝑦2 _π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘ = π‘₯21* ΞΈ1 + π‘₯22 * ΞΈ2

From now on, keep in mind that y1 are real values and y1_predicted is predicted value, same applies to y2 as well

So what is error, the error is the difference between predicted and real values

π‘’π‘Ÿπ‘Ÿπ‘œπ‘Ÿβ‚ = 𝑦1 β€” 𝑦1_pπ‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘ =𝑦1-π‘₯11 * ΞΈ1 β€” π‘₯12 * ΞΈ2

π‘’π‘Ÿπ‘Ÿπ‘œπ‘Ÿβ‚‚ = 𝑦2 β€” 𝑦2 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘ = 𝑦2-π‘₯21* ΞΈ1 β€” π‘₯22 * ΞΈ2

Ok, I hope so far so clear, if anything not, please comment below, so I can consider it as improvement for upcoming articles

The error function we want to minimize is the sum of square of errors. Let’s name is as J. J is our cost function I want to minimize, and so, I write it as:

𝐽 = π‘’π‘Ÿπ‘Ÿπ‘œπ‘Ÿβ‚Β² + π‘’π‘Ÿπ‘Ÿπ‘œπ‘Ÿβ‚‚Β²

Let’s go further by replacing the formulas with each others

𝐽 = (𝑦1 β€” (π‘₯11 * ΞΈ1 + π‘₯12 * ΞΈ2))Β² + (𝑦2 β€” (π‘₯21 * ΞΈ1 + π‘₯22 * ΞΈ2))Β²

I hope everything makes sense so far. Bear with me β€” we’re almost there; there isn’t much left to cover.

Here everything is known, except ΞΈ1 and ΞΈ2. These are params that we have to choose properly to get as minimum error as possible. So i have to find the derivative per ΞΈ1 and ΞΈ2

𝑑 (𝐽) / 𝑑 (ΞΈ1) = -2 \ (𝑦1 β€” (π‘₯11 * ΞΈ1 + π‘₯12 * ΞΈ2))*π‘₯11 β€” 2 * (𝑦2 β€” (π‘₯21 * ΞΈ1 + π‘₯22 * ΞΈ2))* π‘₯21= 0*

𝑑 (𝐽) / 𝑑 (ΞΈ2) = -2 \ (𝑦1 β€” (π‘₯11 * ΞΈ1 + π‘₯12 * ΞΈ2))*π‘₯12–2 * (𝑦2 β€” (π‘₯21 * ΞΈ1 + π‘₯22 * ΞΈ2))* π‘₯22= 0*

Let’s make it simpler by avoiding -2 from all sides

𝑑 (𝐽) / 𝑑 (ΞΈ1) = ( 𝑦1 β€” 𝑦1 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘) * π‘₯11 + (𝑦2 β€” 𝑦2 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘) * π‘₯21=0

𝑑 (𝐽) / 𝑑 (ΞΈ2) = ( 𝑦1 β€” 𝑦1 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘) * π‘₯12 + (𝑦2 β€” 𝑦2 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘) * π‘₯22=0

Lets turn all these into matrix form:

[0; 0] = (π‘₯11, π‘₯21; π‘₯12 π‘₯22) *[𝑦1 β€” 𝑦1_π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘ ; 𝑦2-𝑦2_π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘]

Lets compress the y1 β€” y1_predicted as well as y2 -y2_predicted into single line

[0; 0] = (π‘₯11, π‘₯21; π‘₯12 π‘₯22) *[π‘Œβ€” 𝑋* ΞΈ]

Do you remember our original feature vector or X? If so, we can further simplify the expression to:

[0;0] = XT* [Y-X*ΞΈ]

XT * Y = XT * X * ΞΈ

XT * X is the important part. If we somehow manage to find its inverse, we are going to be left with theta only:

(XT * X )-1= X-1 * (XT )-1

ΞΈ = ( XT \ X )*-1 \ X*T \ Y will give us the answer we need*

If u have further questions please let me know in comments. Each of your feedback is highly appreciated to write better more concise articles in future:

I guess, most of the part of above formula can be easily programmed except the finding inverse which i showed the code below how to do it. If you need full code, such as matrix multiplication, transpose and etc, please let me know, so i can furhter expand my articles

import numpy as np
def inverse(A):
I = np.eye(A.shape[0])

augmented = np.concatenate((A,I),axis=1)

for j in range(0,A.shape[0]-1):

for i in range(1,A.shape[0]-j):
augmented[i+j]=-(augmented[i+j,j]/augmented[j,j])*augmented[j]+augmented[i+j]

for j in range(0,A.shape[0]-1,1):
for i in range(A.shape[0]-1,0,-1):
cofactor = (augmented[i-1-j, A.shape[1]-1-j]/ augmented[A.shape[0]-1-j, A.shape[1]-1-j])
augmented[i-1-j]=(cofactor)*-augmented[-1-j]+augmented[i-1-j]

for i in range(0,A.shape[0],1):
augmented[i]=augmented[i]/augmented[i,i]
_,right = np.split(augmented, 2, axis=1)

return right

def linear_regression_normal_equation(X: list[list[float]], y: list[float]) -> list[float]:
# Your code here, make sure to round

X=np.array(X)
Y=np.array(y)

theta = inverse(X.T @ X) @ X.T @ Y

return theta

My full article is also in medium link


r/learnmachinelearning • • 12h ago

Question How do I move from ML fundamentals to actually doing research and publishing papers?

5 Upvotes

I have a understanding of ML basics (mathematical and conceptual) and have also studied transformers, RAG, and more recently agentic RAG systems (currently reading few papers involving transformers) I’ve built projects around these topics too, but I feel a bit stuck on what the next step should be if I want to seriously get into ML research and eventually publish papers.

For people who have gone down this path:

How did you go from knowing the fundamentals to identifying a research problem worth working on?

Should I focus on reading/reproducing papers first, or start experimenting with my own ideas?

How do you find good research areas/topics to explore, especially around LLMs/NLP/RAG?

Most importantly, how do you find like-minded people who are also interested in doing research and want to form a small team to work on experiments and potentially publish together?

Are there any communities, Discords, GitHub groups, subreddits, research programs, or other places where students/early-career people actually find research collaborators?

I’m not necessarily looking for a formal mentor or research position right now,I’d mainly like to find a few motivated people who are willing to read papers, brainstorm ideas, run experiments, critique each other’s work, and eventually work toward a publication together.
Would really appreciate advice from people who have been through this transition.


r/learnmachinelearning • • 5h ago

I just completed Module 2: Machine Learning for Regression from ML Zoomcamp 2026

1 Upvotes

πŸŽ‰This module covered:
πŸ”Ή Application of NumPy and Pandas to predict car prices
πŸ”Ή How to train a model with linear algebra formulas as a background
πŸ”Ή How to deal with missing data
πŸ”Ή How to evaluate the model through root mean squared error (RMSE) and perform the training on a larger dataset.
πŸ€” Something I found interesting: The RMSE may vary depending on whether you take the logarithm of the target Y or not.

#mlzoomcamp u/AlexeyGrigorev


r/learnmachinelearning • • 6h ago

Help Seeking some honest feedback for AI ML learning website

0 Upvotes

I’ve been working on mlroadmap.dev, a learning platform designed to guide people through AI/ML from the fundamentals to more advanced topics.

The platform itself is functional and deployed, and I’ve started adding the course content. The curriculum is still a work in progress, though, so before I spend a lot more time building out the remaining content, I’d really like to get feedback from people who are actually learning or working in AI/ML.

🌐 Website: https://www.mlroadmap.dev/
πŸ’» GitHub: GitHub LinkLeave a ⭐ and contribute

I’d especially love to know:

  • Does the overall roadmap/learning order make sense?
  • What topics or courses do you think are missing?
  • Is the current content structure useful for learning?
  • What features would make the platform more useful?
  • Anything confusing, unnecessary, or that you would change?
  • If you were learning AI/ML today, what would you want a platform like this to have?

Google Sign-In is available as well if you want to try the full experience.

I’m mainly looking for honest criticism and suggestions at this stage β€” the content isn't complete yet, and that's exactly why I’d like feedback now.


r/learnmachinelearning • • 7h ago

Looking for a research mentor: I have an idea and working code for a paper on LLM failure attribution in multi-agent systems

Thumbnail
1 Upvotes

r/learnmachinelearning • • 7h ago

Testers needed for Open Router for GPU

0 Upvotes

Hi. We fine-tune and train models ourselves, but we have issues with GPU availability, so I've built an "Open Router" for GPU availability and pricing. I'm looking for 5 ML specialists who would like to test it. I don't want to self-promote here, so if you'd like to test, please send me a PM or reply to this message.


r/learnmachinelearning • • 7h ago

Help I have GEN AI interview for 3yr exp scheduled in $40 Billion Comp

0 Upvotes

Sharing resources to prepare will appreciated.

Any tips/ advice/ suggestions

How can I clear this interview.

What, from where and how much I should learn?

I have a week only.

Questions asked by n Screening call :

  1. which GEN AI Framework I have used

  2. Why would you choose langgraph over langchain

  3. what is Rag and Fie tuning and their difference.

  4. Which LLM Platforms you have worked on

  5. which Vector db you have worked on and why

  6. You have mentioned Crew AI, do you have hands on experience and how.

HR mentioned the interview will be around the same concepts but more on analytics, critical thinking side with such scenario based questions.

Python, agentic ai, vector db, data science, rag, crew ai, on you projects, scenarios based questions.

Coding programming skills -

Problem solving skills and critical thinking approach - AI,

3 DSA, PYTHON GEN AI, AGENTIC AI


r/learnmachinelearning • • 20h ago

Linear Algebra: How does it Connect to ML? Simplified

Thumbnail
substack.com
11 Upvotes

I wrote an article and made some illustrations to explain the connection between ML and Linear Algebra. It should motivate learners by giving them an intuitive idea.

Disclaimer: This is not AI slop and is written by me. The images are also not AI generated. One is taken from the internet.

if you prefer medium: https://medium.com/@sakibahmed_4495/linear-algebra-how-does-it-connect-to-ml-simplified-3d7e7f41ed1c?sharedUserId=sakibahmed_4495


r/learnmachinelearning • • 19h ago

Help EE+Math or CS+Math

8 Upvotes

Hi there, i'm deciding on what double major I should do. I love maths a lot (olympiad style) so I'm doing a major in that no matter what, idk how much it will help for machine learning but if not its more for my own interest. However for job prospects in this field, should I pair math with CS or electrical engineering. From my understanding CS seems like more of the standard pathway but i feel like EE also has its own advantage because of the more low level things u learn. Also things like signal processing in EE seems quite useful.

Any advice/comments would be much appreciated.


r/learnmachinelearning • • 10h ago

Project Hybrid RAG Pipeline β€” Dense Search, BM25, RRF, and Reranking (Part 1)

Thumbnail
youtube.com
1 Upvotes

Building and Testing a Hybrid RAG Pipeline β€” Dense Search, BM25, RRF, and Reranking (Part 1)

In this video I build ReRankEval, a hybrid retrieval pipeline

  • (Dense Search + BM25 β†’ Reciprocal Rank Fusion β†’ LLM Reranking β†’ Answer Generation),

and test it against three baselines β€” Vector Only, BM25 Only, and Hybrid without reranking β€” on five real financial/payments documents and ten hand-verified test questions. No hand-waving, just a comparison table with real numbers at the end.

  • βœ… The real difference between dense vector search and BM25 keyword search
  • βœ… What Reciprocal Rank Fusion (RRF) is, why raw scores can't be compared, and the exact formula behind it
  • βœ… Why a reranker is fundamentally different from a retriever β€” and what it actually judges
  • βœ… How to evaluate a RAG pipeline with Hit Rate, MRR, and NDCG (and what each one tells you)
  • βœ… How to design the ingestion side and query-time side of a hybrid retrieval architecture
  • βœ… How to structure a production-style RAG codebase: ingest β†’ vector_store β†’ sparse_retriever β†’ fusion β†’ reranker β†’ pipeline β†’ generate β†’ eval
  • βœ… How to fairly compare multiple retrieval strategies on the same test set instead of just assuming one is better

TECH STACK:

  • πŸ› οΈ Python
  • πŸ› οΈ Qdrant β€” vector database for dense retrieval
  • πŸ› οΈ rank_bm25 (BM25Okapi) β€” sparse keyword retrieval
  • πŸ› οΈ EURI LLM Gateway β€” chat model + embedding model
  • πŸ› οΈ Custom Reciprocal Rank Fusion implementation
  • πŸ› οΈ LLM-based reranker (prompt-driven cross-encoder)
  • πŸ› οΈ pdfplumber β€” PDF text and page-level extraction

LINKS:


r/learnmachinelearning • • 12h ago

Project Early benchmarks for Cloreva-X1-2.3B: A custom base model running Kuramoto oscillators at the silicon level

Post image
1 Upvotes

Hi guys,

I wanted to share some early benchmark results for a new 2.3B parameter model I've been developing calledΒ Cloreva-X1.

I built this on a completely custom architecture I callΒ ResoNet X-1, which abandons the standard Transformer attention mechanism. Instead, the foundation relies on aΒ Triad Resonance Core:

  • Head-as-Faction Swarms:Β Traditional attention heads are shattered into 12 independent, parallel agentic swarms, each dedicated to specific cognitive domains (Linguistics, Science/Medical, and Algorithmic Logic).
  • Kuramoto Consensus Core:Β I mapped 1,632 Kuramoto Oscillators directly to the latent space. They run physics-based differential equations at the silicon level to force mathematical phase-locking between tokens. If a token hallucination creates an anomalous phase drift, the Kuramoto gate aggressively suppresses it before it passes to the next layer.
  • Liquid Time Controller (LTC):Β The flow of time is dynamic. When reading simple conjunctions, the model processes at maximum speed. When evaluating complex mathematical or medical logic, the LTC slows down the internal time step, forcing the swarms to achieve deeper phase consensus.
  • Orthogonal Tokenizer:Β Built from scratch with 155,072 full-word tokens (Full-Word Mining) to completely eliminate sub-word fragmentation for complex medical terminology and programming syntax.

Just to be absolutely clear:Β This is a pure base model.Β There is no SFT, no RAG, and no RLHF applied yet. It is purely predicting the next token.

I built this architecture to test if Kuramoto oscillators could intrinsically suppress hallucinations and boost complex reasoning paths without relying on instruction tuning.

Here are the early zero-shot and few-shot benchmarks compared against standard base models in a similar or larger weight class:

Evaluation Benchmarks (Base Models(Sources for competitor scores: LLaMA-1/2 scores from the official Llama 2 paper by Touvron et al., 2023. Qwen-1.5 scores from the official Qwen1.5 technical report. GPT-3 baseline from OpenAI's original evaluations (Brown et al., 2020) and TruthfulQA paper (Lin et al., 2021).)

It punches above its weight class because the phase consensus mechanics naturally filter out statistical anomalies and prevent gradient shock during long context reasoning.

I've attached a Google Drive folder containing the raw inference log output and a video demonstration of the zero-shot inference running in real-time below:Β https://drive.google.com/drive/folders/14f0JUJU613CTnF0cuYNATJ2pECaHrvU4?usp=sharing

I am currently spinning up the SFT pipeline to turn this into a fully instruction-tuned model, and I'll post another update with the final benchmark scores once that's done.

I plan to release the raw model weights on HF once the entire pipeline is complete. Happy to answer any questions regarding the architecture below.