Skip to main content
SandRise logo SandRise
Exploring Next / By company / Llama 3 1

Topic

Llama 3 1

3 episodes

  1. Ep 402 May 14, 2026

    Many Shot CoT ICL: Making In Context Learning Truly Learn

    Justy and Cody dig into a paper arguing that long-context chain-of-thought prompting behaves less like stuffing a prompt with relevant examples and more like teaching the model during inference. They unpack why many-shot tricks from classification break on reasoning, why semantic retrieval stops helping, and how the paper’s Curvilinear Demonstration Selection tries to order examples like a smooth mini-curriculum.

    TrainingEvalsLlama 3 1Qwen 3
  2. Ep 358 May 1, 2026

    Google AI breakthrough means chatbots use six times less memory during conversations without compromising performance

    Google's TurboQuant compresses AI working memory (the KV cache) by up to 6x in real time using two novel techniques — PolarQuant and QJL — without degrading model performance. Justy and Cody dig into what this actually means for inference costs, who benefits first, and why the 'DeepSeek moment' framing is both apt and a little overblown.

    InferenceLaunchGoogleTurboquant
  3. Ep 212 Mar 10, 2026

    New KV cache compaction technique cuts LLM memory 50x without accuracy loss

    MIT researchers developed Attention Matching, a KV cache compaction technique that achieves 50x memory reduction in LLMs without accuracy loss, solving a critical bottleneck for enterprise applications handling long contexts.

    InferenceMitLlama 3 1Qwen 3
SandRise logo SandRise Product Studio
Resume LinkedIn GitHub Email

© 2026 SandRise · Built by Nick Sanders

🧠 PM Perspective

Crafting your PM challenge
Analyzing context and generating a thoughtful question...
Your Challenge
0 / 2000
✨

Feedback on Your Answer

⚠️

Say Hi

Feedback, ideas, interesting finds — anything goes.

What's this about?
0 / 2,000

Note received!

Thanks for reaching out. I'll take a look soon.

⚠️

Something went wrong. Please try again.