Based on recent Ornith and deepseek results, this research was outdated before publication.
This space moves too fast for any more traditional scientific research to actually keep up with anymore imo.
The goal post for that “trillion dollar question” keeps moving. At what point do people realize that even current models are capable of self improvement in the right test setup (Ornith, Qwen), or that fitting the right info into the context window (skills files in general, Hermes Agent’s own context-only self improvement loop, etc), or that creating multiple specialized models followed by distillation (deepseek v4’s teacher models) all, in fact, work quite well?
Based on recent Ornith and deepseek results, this research was outdated before publication.
This space moves too fast for any more traditional scientific research to actually keep up with anymore imo.
The goal post for that “trillion dollar question” keeps moving. At what point do people realize that even current models are capable of self improvement in the right test setup (Ornith, Qwen), or that fitting the right info into the context window (skills files in general, Hermes Agent’s own context-only self improvement loop, etc), or that creating multiple specialized models followed by distillation (deepseek v4’s teacher models) all, in fact, work quite well?