Skip to main content
SandRise logo SandRise
Exploring Next / By company / Gpt 2

Topic

Gpt 2

3 episodes

  1. Ep 445 Jun 1, 2026

    Serving Multiple Users at Once: How Continuous Batching Keeps LLM Inference Efficient MachineLearningMastery

    Continuous batching is a scheduling technique that keeps LLM inference servers from wasting GPU cycles on padding. Instead of forcing short requests to wait for long ones in a fixed batch, continuous batching frees up slots the moment a request finishes and admits new work immediately, eliminating idle padding tokens and improving throughput.

    InferenceGpt 2Blog
  2. Ep 309 Apr 20, 2026

    6 Things I Learned Building LLMs From Scratch That No Tutorial Teaches You | Towards Data Science

    Justy and Cody dig into what actually changes when you stop calling an LLM API and start building pieces yourself: why fine-tuning tricks like RsLoRA matter, why RoPE won, where weight tying still makes sense, why Pre-LN became the default, and how KV cache buys speed by spending memory.

    TrainingInferenceGpt 2Llama
  3. Ep 31 Nov 21, 2025

    Let’s Build the GPT Tokenizer: A Complete Guide to Tokenization in LLMs – fast

    18 months ago, Andrej Karpathy set a challenge : “Can you take my 2h13m tokenizer video and translate the video into the format of a book chapter”. We’ve done it, and the chapter is below, including key pieces of code inlined, and images from the video at key points (hyperlinked to the video timestamp).

    Dev ToolsFast AIAndrej KarpathyGpt 2
SandRise logo SandRise Product Studio
Resume LinkedIn GitHub Email

© 2026 SandRise · Built by Nick Sanders

🧠 PM Perspective

Crafting your PM challenge
Analyzing context and generating a thoughtful question...
Your Challenge
0 / 2000
✨

Feedback on Your Answer

⚠️

Say Hi

Feedback, ideas, interesting finds — anything goes.

What's this about?
0 / 2,000

Note received!

Thanks for reaching out. I'll take a look soon.

⚠️

Something went wrong. Please try again.