AI That "Understands" Much Later: The Mystery of Grokking
Imagine: You train a neural network. It quickly memorizes answers on training examples, but on new data it shows complete helplessness. Then, after a very long time of training, it suddenly "understands" the task and starts giving correct answers. This phenomenon is called Grokking (from the jargon meaning "to deeply understand the essence"). 📌 The Core Discovery: A team of researchers (Mingyue Xu, Gal Vardi, Itay Safran) has rigorously proven this strange behavior in a simple linear regression model and showed how to control it. 🎯 What is the Problem? Usually, we think that the longer an AI trains, the better it generalizes. But in the case of Grokking, the opposite happens: ① Overfitting → ② Long Plateau → ③ Sudden "Understanding" That is, the model first memorizes the data (overfits), then stagnates for a long time , and then suddenly "understands" the underlying pattern. 📊 What Did the Scientists Prove? For the...