Play social multiplayer games against frontier AI models and see if you can beat them! [D]
play here Play games like poker, risk, diplomacy with friends or alone against AI models, guess whatβ¦
AI tools, cybersecurity and development news aggregated from top sources β saved permanently with unique URLs.
Play social multiplayer games against frontier AI models and see if you can beat them! [D]
play here Play games like poker, risk, diplomacy with friends or alone against AI models, guess whatβ¦
Simulating fault tolerance with stage skipping in pipeline-parallel training [R]
Our most recent work at Templar explores fault tolerance in Crucible, our distributed pre-training pβ¦
LinearSolveBench: new benchmark for linear solvers [P]
LinearSolverBench measures the ability of a model or harness to write fast, accurate, and general nuβ¦
QontoFAQ: A better Information Retrieval Benchmark [R]
Retrieval benchmarks sometimes feel benchmaxxed by models, so we wanted to find a way to tie it as cβ¦
Understanding and Enhancing Kimi Delta Attention [R]
TLDR: We demonstrate and explain the difference in expressivity of Gated Deltanet (GDN) and Kimi Delβ¦
OpenTrainDNN: A Browser-Based Real-Time Neural Network Visualizer. [P]
OpenTrainDNN is an open-source, client-side web application designed to render the step-by-step traiβ¦
Xiaomi releases MiMo-V2.6: "Frontier intelligence, all the modalities, built in public." [N]
The total training cost was just $3.5M. The model comes with a live benchmaxxing dashboard. https://β¦
Paper on ArXiv for a year now, should I disclose about it in ICLR submission? [Discussion]
Hi, I have a solo paper on arxiv from my masters studies, which was my part-time work and a few montβ¦
How bad is it if training and validation data overlap in a student ML project?[D]
Iβm working on a student machine learning/computer vision project and recently realized that my valiβ¦
Jev's calibration was measured. The LLMs won [D]
Source: Jev Benchmarks Its training method is literally called "Reinforcement Learning for Calibrateβ¦
I built a framework-free prototype learner that lets local LLMs learn and correct facts instantly (1.6xβ4x faster than backprop)[R]
Hey everyone, I wanted to share a project Iβve been working on called Jayce. The whole thing startedβ¦
Jev vs Laya head to head benchmark [D]
I asked Fable 5.1 to benchmark Jev vs Laya on accuracy and speed. Here's the full report: https://clβ¦