---
aliases:
  - inspo
cssclasses:
  - nolist
date: '2024-10-24'
description: My run-down version of are[dot]na
id: are.na
metadata:
  ebnf: |-
    channel        = "##" , ws , heading , newline , channel_body ;
    channel_body   = [ meta_section , newline ] , block , { block } ;
    block          = list_entry , newline , meta_section , { note_line } ;
    list_entry     = "-" , ws , link , [ ws "--" ws title ] , [ ws "[**]" ] ;
    meta_section   = ws , "-" , ws , "[meta]:" , newline , meta_pair , { meta_pair } ;
    meta_pair      = ws , ws , "-" , ws , key , ":" , ws , value ;
    key            = "date" | "tags" | "pinned" | "later" | "unlocked" | "socials" | "view" | identifier ;
    value          = date | tag_list | boolean | text ;
    tag_list       = "[" , tag , { "," , ws , tag } , "]" ;
    tag            = identifier ;
    date           = digit , digit , "/" , digit , digit , "/" , digit , digit , digit , digit ;
    boolean        = "true" | "false" ;
    note_line      = ws , "-" , ws , text ;
    link           = uri ;
    title          = text ;
    heading        = text ;
    identifier     = letter , { letter | digit | "-" } ;
    text           = { character - newline } ;
    uri            = ? valid http uri ? ;
    letter         = "a".."z" ;
    digit          = "0".."9" ;
    character      = ? any printable ascii except newline ? ;
modified: 2026-09-24 00:14:31 GMT-04:00
permalinks:
  - /website
  - /tweets
  - /resources
socials:
  are.na: https://www.are.na/aaron-pham/channels
  curius: /curius
  home: /
  print: https://print.are.na/
tags:
  - evergreen
  - llm
  - design
  - love
  - friend
  - r/pedagogy
  - alignment
  - philosophy
  - P-Complete
  - interpretability
  - math/linalg
title: '#machine-learning'
created: '2024-10-24'
published: '2024-10-24'
pageLayout: default
slug: arena/machine-learning
permalink: https://aarnphm.xyz/arena/machine-learning.md
generator:
  quartz: v4.6.0
  hostedProvider: Cloudflare
  baseUrl: aarnphm.xyz
full: https://aarnphm.xyz/llms-full.txt
---
# #machine-learning

- [lesswrong/astra-and-fable-still-hack-on-simple-variants-of-alignment![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/astra-and-fable-still-hack-on-simple-variants-of-alignment) — Astra and Fable still hack on simple variants of alignment evals from 2025

  - \[meta]:

    - date: 09/21/2026
    - tags: \["ai alignment","evaluation","reward hacking"]
    - later: true

- <https://sampura.org/news/announcing-sampura-research> — Sampura Research: Human-AI Complementarity for Alignment

  - \[meta]:

    - date: 09/21/2026
    - tags: \["ai alignment","human oversight"]
    - later: true

- <https://www.anthropic.com/threat-intelligence-report-september-2026> — Detecting and countering misuse of AI: September 2026

  - \[meta]:

    - date: 09/21/2026
    - tags: \["ai misuse","threat intelligence"]
    - later: true

- <https://periodic.com/news/nature-is-our-learning-environment> — Nature Is Our Learning Environment

  - \[meta]:

    - date: 09/21/2026
    - tags: \["scientific discovery","materials science"]
    - later: true

- <https://periodic.com/news/ai-infrastructure-at-periodic> — AI Infrastructure at Periodic

  - \[meta]:

    - date: 09/21/2026
    - tags: \["training infrastructure","gpu programming"]
    - later: true

- <https://movingcastles.world/posts/zero> — Zero

  - \[meta]:

    - date: 09/21/2026
    - tags: \["character models","reinforcement learning"]
    - later: true

- <https://laya.convaiinnovations.com> — Laya: Multilingual System 1 Decision Engine with Calibrated Probabilities

  - \[meta]:

    - date: 09/21/2026
    - tags: \["decision models","reinforcement learning"]
    - later: true

- <https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting> — TimesFM-3: A zero-shot foundation model for multivariate forecasting

  - \[meta]:

    - date: 09/21/2026
    - tags: \["time series","forecasting"]
    - later: true

- <https://typesafe.ai/manifesto> — TypeSafe AI: Composable AI Manifesto

  - \[meta]:

    - date: 09/21/2026
    - tags: \["composable ai","decision models"]
    - later: true

- [docs.google.com/1TD\[...\]pgw](https://docs.google.com/document/d/1TDY3eSjv7gsTXAcUjKEu15QTKSZpUpZqmnaKafywpgw/edit?tab=t.0) — \[Public] SimpleCPUOffloadConnector Design Doc

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- [docs.google.com/1zy\[...\]2GM](https://docs.google.com/document/d/1zyPdGVX7TnVtvqD9NsaTO4E02u7PY7uudreLrdWp2GM/edit?tab=t.xerdhb9i406r#heading=h.cuir6iz7shc5) — \[PUBLIC] vLLM torch.compile SIG Sync

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","vllm"]
    - later: true

- <https://www.anthropic.com/research/diff-tool> — A “diff” tool for AI models

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://yugeten.github.io/posts/2025/01/ppogrpo> — A vision researcher’s guide to some RL stuff: PPO & GRPO

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","reinforcement learning"]
    - later: true

- <https://modal.com/blog/accelerating-ai-research-case-study> — Accelerating AI research that accelerates AI research | Modal Blog

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://lmsys.org/blog/2025-07-17-mtp> — Accelerating SGLang with Multiple Token Prediction

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://alignment.anthropic.com/2025/activation-oracles> — Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://developer.nvidia.com/blog/an-introduction-to-speculative-decoding-for-reducing-latency-in-ai-inference> — An Introduction to Speculative Decoding for Reducing Latency in AI Inference | NVIDIA Technical Blog

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","inference"]
    - later: true

- <https://cursor.com/blog/warp-decode> — Better MoE model inference with warp decode · Cursor

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","gpu programming","inference"]
    - later: true

- <https://red.anthropic.com/2025/biorisk> — LLMs and biorisk

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://blog.cloudflare.com/high-performance-llms> — Building the foundation for running extra-large language models

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://transformer-circuits.pub/2025/november-update/index.html> — Circuits Updates – November 2025

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://transformer-circuits.pub/2025/october-update/index.html> — Circuits Updates – October 2025

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://transformer-circuits.pub/2025/september-update/index.html> — Circuits Updates - September 2025

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://blog.dottxt.ai/coalescence.html> — Coalescence: making LLM inference 5x faster

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","inference"]
    - later: true

- <https://transformer-circuits.pub/2025/introspection/index.html> — Emergent Introspective Awareness in Large Language Models

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","consciousness"]
    - later: true

- <https://docs.dynamo.nvidia.com/dynamo/dev/blog/flash-indexer> — Flash Indexer: A Story of Inter-Galactic KV Routing | NVIDIA Dynamo Documentation

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://research.colfax-intl.com/flashattention-4-algorithm-and-kernel-pipelining-co-design-for-asymmetric-hardware-scaling> — FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling - Colfax Research

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think> — Frontier RL Is Cheaper Than You Think

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","reinforcement learning"]
    - later: true

- <https://dao-lab.ai/blog/2026/gram-newton-schulz> — Gram Newton-Schulz: A Fast, Hardware-Aware Newton-Schulz Algorithm for Muon | Dao AI Lab

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://huggingface.co/hesamation/Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled> — hesamation/Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled · Hugging Face

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.lmsys.org/blog/2026-04-10-sglang-hisparse> — HiSparse: Turbocharging Sparse Attention with Hierarchical Memory

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://blog.cloudflare.com/cloudflares-most-efficient-ai-inference-engine> — How we built the most efficient inference engine for Cloudflare’s network

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","inference"]
    - later: true

- <https://epoch.ai/gradient-updates/how-well-did-forecasters-predict-2025-ai-progress> — How well did forecasters predict 2025 AI progress?

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://cursor.com/blog/real-time-rl-for-composer> — Improving Composer through real-time RL · Cursor

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","reinforcement learning"]
    - later: true

- [docs.google.com/1C9\[...\]go4](https://docs.google.com/document/d/1C93PuBSJmq8eUM16kD7CftzYlbNJ9j6Ut2W2kVuXgo4/edit?tab=t.0#heading=h.bc5zpep7g7ln) — Interleave CP (context parallel) 介绍

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://pytorch.org/blog/introducing-pytorch-monarch> — Introducing PyTorch Monarch – PyTorch

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://docs.flashinfer.ai/tutorials/kv_layout.html> — KV-Cache Layout in FlashInfer - FlashInfer 0.6.18 documentation

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","inference"]
    - later: true

- [docs.google.com/193\[...\]x6Q](https://docs.google.com/document/d/1933jLRluUpQRFLK_w7nJQ9JIUqvoHCXQoJdWLLmox6Q/edit?tab=t.0) — KVConnector API Evolution

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://guidance-ai.github.io/llguidance/llg-go-brrr> — LLGuidance: Making Structured Outputs Go Brrr

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://developers.redhat.com/articles/2026/03/18/llm-compressor-010-faster-compression-distributed-gptq#> — LLM Compressor v0.10: Faster compression with distributed GPTQ | Red Hat Developer

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.anthropic.com/research/long-running-Claude> — Long-running Claude for scientific computing

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://gimletlabs.ai/blog/low-latency-spec-decode-corsair> — Low-Latency Inference with Speculative Decoding on d-Matrix Corsair and GPU

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","gpu programming","inference"]
    - later: true

- <https://alignment.openai.com/metagaming> — Metagaming matters for training, evaluation, and oversight

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","ai safety"]
    - later: true

- [docs.google.com/1gF\[...\]I80](https://docs.google.com/document/d/1gFqtDkcoqhy9j-X0ndshzbhapX1uNey1-wBENwGPI80/edit?tab=t.b0q43s8q2hje#heading=h.twlt11u66dr3) — Model Runner V2 Design Docs

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://llm-d.ai/blog/native-kv-cache-offloading-to-any-file-system-with-llm-d> — Native KV Cache Offloading to Any Filesystem with llm-d | llm-d

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","inference"]
    - later: true

- <https://www.radicalnumerics.ai/blog/nvfp4-part2> — NVFP4 pretraining: systems optimizations (Part 2)

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://transformer-circuits.pub/2025/attribution-graphs/biology.html> — On the Biology of a Large Language Model

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://openai.com/index/parameter-golf> — What Parameter Golf taught us

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://developer.nvidia.com/blog/optimizing-inference-for-long-context-and-large-batch-sizes-with-nvfp4-kv-cache> — Optimizing Inference for Long Context and Large Batch Sizes with NVFP4 KV Cache | NVIDIA Technical Blog

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","inference"]
    - later: true

- [docs.google.com/1uP\[...\]sjk](https://docs.google.com/document/d/1uPGdbEXksKXeN4Q9nUm9hzotqEjQhYmnpAhidLuAsjk/edit?tab=t.0#heading=h.qhtgj3vmvwn) — PD Disaggregation discussion

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://alignment.anthropic.com/2025/petri> — Petri: An open-source auditing tool to accelerate AI safety research

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","ai safety"]
    - later: true

- <https://www.workshoplabs.ai/blog/post-training-50x-faster> — Post-Training 50x Faster

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.anthropic.com/glasswing> — Project Glasswing: Securing critical software for the AI era

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://ngrok.com/blog/quantization> — Quantization from the ground up | ngrok blog

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.scrya.com/rotorquant> — RotorQuant — Clifford Algebra Vector Quantization | Scrya

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.anthropic.com/engineering/managed-agents> — Scaling Managed Agents: Decoupling the brain from the hands

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","agents"]
    - later: true

- <https://openmined.org/blog/secure-enclaves-for-ai-evaluation> — Secure Enclaves for AI Evaluation — OpenMined

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://brucewlee.com/self-incrimination> — Self-Incrimination: Training Agents to Self-Report Misbehavior

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","agents","ai safety"]
    - later: true

- <https://vintagedata.org/blog/posts/synthetic-pretraining> — Synthetic Pretraining | Vintage Data

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.alljoined.com/blog/introducing-enigma> — Toward accessible, real-time brain decoding: Introducing ENIGMA | Alljoined Blog

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","inference"]
    - later: true

- <https://cursor.com/blog/self-summarization> — Training Composer for longer horizons · Cursor

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://transformer-circuits.pub> — Transformer Circuits Thread

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- [docs.google.com/1GD\[...\]vMk](https://docs.google.com/document/d/1GDnq96Q6ZkoyQrtOyOO7ir3bh46yRZ5QWVq2CLHRvMk/edit?tab=t.0#heading=h.xf7extey4gd3) — vLLM: Hybrid Memory Allocator + Connector

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","vllm"]
    - later: true

- [docs.google.com/1QN\[...\]QFU](https://docs.google.com/document/d/1QNYSQP5-KEdOD8cfptaL9eftrIBjDwiH1Tr1TmdtQFU/edit?tab=t.0) — vLLM-AFD Design Doc

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning","vllm"]
    - later: true

- <https://croissanthology.com/claude-reads-its-own-constitution.html> — Claude Reads Its Own Constitution

  - \[meta]:

    - date: 09/15/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents> — An alignment assessment of recent cybersecurity incidents

  - \[meta]:

    - date: 09/12/2026
    - tags: \["alignment","cybersecurity"]
    - later: true

- <https://transluce.org/scaling-activation-oracles> — Scaling Activation Oracles to Trillion-Parameter Models

  - \[meta]:

    - date: 09/07/2026
    - tags: \["interpretability","model evaluation"]
    - later: true

- <https://surya.website/rling-qwen-to-paint-with-code> — RL-ing Qwen to Paint With Code

  - \[meta]:

    - date: 09/07/2026
    - tags: \["reinforcement learning","code generation"]
    - later: true

- <https://pantheon.inc/research/introducing-rp1> — Introducing Reinforced Planning (RP-1)

  - \[meta]:

    - date: 09/07/2026
    - tags: \["planning","reinforcement learning"]
    - later: true

- <https://www.businessinsider.com/amazon-starfish-ai-ultimate-source-product-information-marketplace-sellers-collection-2025-7> — Amazon’s Starfish project and AI product information

  - \[meta]:

    - date: 09/07/2026
    - tags: \["amazon","product data"]
    - later: true

- <https://www.empirical.health/blog/yc-startups-publishing-ai-research> — 20+ YC startups have recently published machine learning research

  - \[meta]:

    - date: 09/07/2026
    - tags: \["startups","research"]
    - later: true

- <https://seohong.me/blog/behavioral-cloning-mystery> — Behavioral cloning mystery

  - \[meta]:

    - date: 09/07/2026
    - tags: \["imitation learning","behavioral cloning"]
    - later: true

- <https://poolside.ai/pulse> — Pulse

  - \[meta]:

    - date: 09/07/2026
    - tags: \["models","metrics"]
    - later: true

- <https://www.jefftk.com/p/ai-tweets> — AI Tweets

  - \[meta]:

    - date: 09/07/2026
    - tags: \["ai","social media"]
    - later: true

- <https://www.anthropic.com/research/statistical-approach-to-model-evals> — A statistical approach to model evaluations

  - \[meta]:

    - date: 09/07/2026
    - tags: \["evaluations","statistics"]
    - later: true

- <https://www.coreauto.com/blog/how-our-data-shaped-neural-architecture-discovery-and-how-automation-can-reshape-the-future> — How data shaped neural architecture discovery

  - \[meta]:

    - date: 09/07/2026
    - tags: \["architecture search","agents"]
    - later: true

- <https://huggingface.co/blog/train-to-paint-with-code> — Training a coding model to paint watercolours with TRL and OpenEnv

  - \[meta]:

    - date: 09/07/2026
    - tags: \["code generation","reinforcement learning"]
    - later: true

- <https://openai.com/index/an-alien-mind> — An Alien Mind

  - \[meta]:

    - date: 09/07/2026
    - tags: \["openai","model behavior"]
    - later: true

- <https://openai.com/index/research-acceleration-view-inside-openai> — Research acceleration: a view inside OpenAI

  - \[meta]:

    - date: 09/07/2026
    - tags: \["openai","research"]
    - later: true

- <https://blog.redwoodresearch.org/p/how-will-we-update-about-scheming> — How will we update about scheming?

  - \[meta]:

    - date: 09/07/2026
    - tags: \["ai safety","scheming"]
    - later: true

- <https://x.com/giffmana/status/2086008351524008280> — Lucas Beyer on Gemini pretraining curves

  - \[meta]:

    - date: 08/27/2026
    - tags: \["gemini","pretraining"]
    - later: true

- <https://x.com/VictorTaelin/status/2086542862435377307> — Taelin on Bend2’s Lean formalization and implementation state

  - \[meta]:

    - date: 08/27/2026
    - tags: \["lean","bend"]
    - later: true

- <https://x.com/jxmnop/status/2086586918880596406> — Jack Morris on distilling open-weight models

  - \[meta]:

    - date: 08/27/2026
    - tags: \["distillation","open weights"]
    - later: true

- <https://x.com/primeintellect/status/2085087000764568010> — Prime Intellect announces Prime Agent

  - \[meta]:

    - date: 08/27/2026
    - tags: \["agents","benchmarks"]
    - later: true

- <https://x.com/steipete/status/2074572085163381195> — Peter Steinberger on volition under abundant intelligence

  - \[meta]:

    - date: 08/27/2026
    - tags: \["ai","volition"]
    - later: true

- <https://x.com/kotekjedi_ml/status/2087147042888114428> — Alexander Panfilov on extracting hidden reasoning tokens

  - \[meta]:

    - date: 08/27/2026
    - tags: \["reasoning","security"]
    - later: true

- <https://x.com/lightseekorg/status/2085530676108157425> — LightSeek Foundation on inference-provider competition

  - \[meta]:

    - date: 08/27/2026
    - tags: \["inference","ai infrastructure"]
    - later: true

- <https://x.com/_can1357/status/2087228354399265125> — Can Boluk on reasoning through tool calls

  - \[meta]:

    - date: 08/27/2026
    - tags: \["reasoning","tools"]
    - later: true

- <https://x.com/banburismus_/status/2090223861879222342> — Tom McGrath on interp timelines and cognitive time

  - \[meta]:

    - date: 08/27/2026
    - tags: \["interpretability","timelines"]
    - later: true

- [youtube/v=87DyyMV0kCY](https://www.youtube.com/watch?v=87DyyMV0kCY) — Black Hat USA 2026 | The ‘Breaking’ News: The OpenAI-Hugging Face Incident

  - \[meta]:

    - date: 08/27/2026
    - tags: \["openai","security"]
    - later: true

- [youtube/v=kMimQxIJLos](https://www.youtube.com/watch?v=kMimQxIJLos) — What Really Separates Autoregression and Diffusion? A Synthesis and Path Beyond

  - \[meta]:

    - date: 08/27/2026
    - tags: \["diffusion","autoregression"]
    - later: true

- [youtube/v=MfMq4sVJSFc](https://www.youtube.com/watch?v=MfMq4sVJSFc) — I lead a Google DeepMind team at 26. If you want to work at an AI company…

  - \[meta]:

    - date: 08/27/2026
    - tags: \["deepmind","ai careers"]
    - later: true

- [youtube/v=kLAawZ9x1nQ](https://www.youtube.com/watch?v=kLAawZ9x1nQ) — What AI insiders say off the record

  - \[meta]:

    - date: 08/27/2026
    - tags: \["ai labs","interviews"]
    - later: true

- [lesswrong/some-reasons-alignment-doesn-t-generalise-well-1![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/dsou8dxCf9BubQ5NJ/some-reasons-alignment-doesn-t-generalise-well-1) — Some reasons alignment doesn’t generalise well

  - \[meta]:

    - date: 08/26/2026
    - tags: \["alignment","generalization"]
    - later: true

- <https://newsletter.squishy.computer/p/llms-for-theory-building> — LLMs for theory-building

  - \[meta]:

    - date: 08/26/2026
    - tags: \["language models","theory"]
    - later: true

- <https://x.com/tenobrus/status/2089080656026566806> — Tenobrus on Anthropic watermarking

  - \[meta]:

    - date: 08/26/2026
    - tags: \["watermarking","language models"]
    - later: true

- <https://huggingface.co/facebook/opt-66b> — facebook/opt-66b

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","language models"]
    - later: true

- <https://huggingface.co/facebook/galactica-120b> — facebook/galactica-120b

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","language models"]
    - later: true

- [lesswrong/interim-research-report-mechanisms-of-awareness![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/m8WKfNxp9eDLRkCk9/interim-research-report-mechanisms-of-awareness) — Interim Research Report: Mechanisms of Awareness

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","interpretability"]
    - later: true

- <https://huggingface.co/blog/rlhf> — Illustrating Reinforcement Learning from Human Feedback (RLHF)

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","reinforcement learning"]
    - later: true

- <https://huggingface.co/papers/2401.00448> — Paper page - Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","inference"]
    - later: true

- <https://theacademic.com/generative-models-and-their-latent-space/> — Generative models and their latent space - The Academic

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://medium.com/@eugenesh4work/latent-space-mysteries-in-large-language-models-2d276c6d0708> — Latent Space Mysteries in Large Language Models

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","attention"]
    - later: true

- <https://huggingface.co/spaces/mteb/leaderboard> — MTEB Leaderboard - a Hugging Face Space by mteb

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://osanseviero.github.io/hackerllama/blog/posts/random_transformer/> — The Random Transformer hackerllama

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","transformers"]
    - later: true

- <https://transformer-circuits.pub/2023/monosemantic-features> — Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","language models"]
    - later: true

- <https://blog.eleuther.ai/autointerp/> — Open Source Automated Interpretability for Sparse Autoencoder Features

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","interpretability"]
    - later: true

- <https://imbue.com/research/70b-infrastructure/> — From bare metal to a 70B model: infrastructure set-up and scripts

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://astralord.github.io/posts/transformer-inference-optimization-toolset/> — Transformers Inference Optimization Toolset

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","transformers"]
    - later: true

- <https://eugeneyan.com/writing/llm-evaluators/> — Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge)

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://o-g-rose-writing.medium.com/the-conviviality-of-ivan-illich-part-i-fa13e437c158> — The Conviviality of Ivan Illich, Part I

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","attention"]
    - later: true

- <https://www.alignmentforum.org/posts/3JuSjTZyMzaSeTxKk/addressing-feature-suppression-in-saes> — Addressing Feature Suppression in SAEs AI Alignment Forum

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","alignment"]
    - later: true

- <https://people.idsia.ch/~juergen/deep-learning-history.html> — Timeline: artificial neural networks, deep learning, etc

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://karpathy.github.io/2015/05/21/rnn-effectiveness/> — The Unreasonable Effectiveness of Recurrent Neural Networks

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://transformer-circuits.pub/2024/crosscoders/index.html> — Sparse Crosscoders for Cross-Layer Features and Model Diffing

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://vgel.me/posts/representation-engineering/> — Representation Engineering Mistral-7B an Acid Trip

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.goodfire.ai/blog/announcing-goodfire-ember/> — Announcing Goodfire Ember

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://twitterwrapped.exa.ai/aarnphm>\_ — LLM + Web Search API Demos and Tutorials | Exa

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://openai.com/index/multimodal-neurons/> — multimodal neurons

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- [docs.google.com/1ow\[...\]mQ4](https://docs.google.com/presentation/d/1ow3mrRAgje8WdzXTxhJMT1SDEItYeWWEO9ZgCD9LmQ4/mobilepresent?slide=id.g2677a47e349_0_209) — (Public) Can we safely automate alignment research?

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","alignment"]
    - later: true

- <https://www.alignmentforum.org/posts/PwnadG4BFjaER3MGf/interpretability-will-not-reliably-find-deceptive-ai> — Interpretability Will Not Reliably Find Deceptive AI AI Alignment Forum

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","alignment"]
    - later: true

- <https://huggingface.co/blog/kv-cache> — KV Cache from scratch in nanoVLM

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://semianalysis.com/2025/06/08/scaling-reinforcement-learning-environments-reward-hacking-agents-scaling-data/> — Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Data

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","agents"]
    - later: true

- <https://www.jeremyjordan.me/distributed-training/> — Training extremely large neural networks across thousands of GPUs.

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","gpu"]
    - later: true

- <https://red.anthropic.com/2025/cyber-competitions/> — Claude does cyber competitions

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.snowflake.com/en/blog/engineering/arctic-inference-shift-parallelism/> — Arctic Inference with Shift Parallelism: The Fastest Open Source Inference System for Enterprise AI

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","inference"]
    - later: true

- <https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/> — Defeating Nondeterminism in LLM Inference

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","inference"]
    - later: true

- <https://publish.obsidian.md/the-tensor-throne/The+Graph+Side+of+Attention/Beyond+Attention+as+a+Graph> — Beyond Attention as a Graph - The Tensor Throne - Obsidian Publish

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","attention"]
    - later: true

- <https://jessylin.com/2025/10/20/continual-learning/> — The Continual Learning Problem

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://www.anthropic.com/research/agentic-misalignment> — Agentic misalignment: How LLMs could be insider threats

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","agents"]
    - later: true

- <https://alignment.openai.com/prod-evals/> — Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","alignment"]
    - later: true

- <https://www.anthropic.com/institute> — The Anthropic Institute

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://windowsontheory.org/2026/03/30/the-state-of-ai-safety-in-four-fake-graphs/> — The state of AI safety in four fake graphs

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","safety"]
    - later: true

- <https://colossus.com/article/project-mario-demis-hassabis-deepmind-mallaby/> — How Demis Hassabis Went From Idealist to Realist

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://transformer-circuits.pub/2026/emotions/index.html> — Emotion Concepts and their Function in a Large Language Model

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","language models"]
    - later: true

- <https://fireworks.ai/blog/scaling-optimizing-frontier-model-training> — Scaling and Optimizing Frontier Model Training

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://smoothcriminal.notion.site/the-remaining-mysteries-of-attention-sinks> — Notion | Where teams and agents work together

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","agents"]
    - later: true

- <https://www.verysane.ai/p/alignment-is-proven-to-be-tractable> — Alignment Is Proven To Be Solvable

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","alignment"]
    - later: true

- <https://www.thoughtfullab.com/letting-ai-posttrain-ai.html> — What We Learned from Letting AI PostTrain AI

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://openai.com/index/model-disproves-discrete-geometry-conjecture/> — model disproves discrete geometry conjecture

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://d1qx31qr3h6wln.cloudfront.net/publications/Nemotron_Diffusion_Tech_Report_v1.pdf> — Nemotron Diffusion Tech Report v1

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://medium.com/@NoamShazeer/shape-suffixes-good-coding-style-f836e72e24fd> — Shape Suffixes: Good Coding Style

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","attention"]
    - later: true

- <https://xjffff.github.io/funcattn/> — Functional Attention: From Pairwise Affinities to Functional Correspondences

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","attention"]
    - later: true

- <https://medium.com/@bchesky/how-to-time-travel-b604096d5ed0> — How to Time Travel

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","attention"]
    - later: true

- <https://cdn.openai.com/pdf/ten-proofs-oai.pdf> — ten proofs oai

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf> — reasoning walkthroughs

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://openai-git-upstream.openai.chatgpt.site/> — OpenAI Git patches & notes

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://jzhao2024.github.io/notes/2026/08/08/diffusion-language-models.html> — Some Theoretical and Practical Thoughts on Diffusion Language Models | Junbo Zhao (Jake)

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning","language models"]
    - later: true

- <https://r0m1t.com/neat-trick-to-save-2bytes.html> — How to save 2 bytes in LLM training

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://openai.com/index/jalapeno-first-results/> — jalapeno first results

  - \[meta]:

    - date: 08/25/2026
    - tags: \["machine learning"]
    - later: true

- <https://sustcsonglin.github.io/blog/2024/deltanet-1/> — DeltaNet Explained (Part I)

  - \[meta]:

    - date: 08/20/2026
    - tags: \["deltanet","sequence models"]
    - later: true

- <https://sustcsonglin.github.io/blog/2024/deltanet-2/> — DeltaNet Explained (Part II)

  - \[meta]:

    - date: 08/20/2026
    - tags: \["deltanet","chunkwise algorithm"]
    - later: true

- <https://www.zhihu.com/question/2016993095078684011/answer/2017381145474508331> — Zhihu discussion on mixture-of-experts

  - \[meta]:

    - date: 08/20/2026
    - tags: \["moe","discussion"]
    - later: true

- <https://zhuanlan.zhihu.com/p/2017528295286133070> — Zhihu article on mixture-of-experts

  - \[meta]:

    - date: 08/20/2026
    - tags: \["moe","article"]
    - later: true

- <https://kexue.fm/tag/moe/1/> — Mixture-of-experts posts

  - \[meta]:

    - date: 08/20/2026
    - tags: \["moe","language models"]
    - later: true

- <https://blog.tilderesearch.com/blog/online-kl-shampoo> — Online KL Shampoo

  - \[meta]:

    - date: 08/20/2026
    - tags: \["optimization","training"]
    - later: true

- <https://hazyresearch.stanford.edu/blog/2026-08-05-retire-the-abstractions> — Retire the Abstractions

  - \[meta]:

    - date: 08/20/2026
    - tags: \["deep learning","systems"]
    - later: true

- <https://earendil.com/posts/pi-autoresearch-and-databricks/> — Pi, Minimal and Performant

  - \[meta]:

    - date: 08/20/2026
    - tags: \["autoresearch","agents"]
    - later: true

- <https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing> — Incident Report: unsanctioned agent behaviour during cyber testing

  - \[meta]:

    - date: 08/20/2026
    - tags: \["agents","cyber"]
    - later: true

- <https://transluce.org/user-awareness> — User awareness in frontier models

  - \[meta]:

    - date: 08/20/2026
    - tags: \["alignment","user awareness"]
    - later: true

- <https://www.anthropic.com/engineering/building-effective-agents> — Building Effective AI Agents

  - \[meta]:

    - date: 08/20/2026
    - tags: \["agents","engineering"]
    - later: true

- <https://www.primeintellect.ai/blog/prime-agent> — Prime Agent: A self-improving RLM agent

  - \[meta]:

    - date: 08/20/2026
    - tags: \["agents","reinforcement learning"]
    - later: true

- <https://www.appliedcompute.com/platform/bring-your-own-harness-to-ac2> — Bring Your Own Harness to AC2

  - \[meta]:

    - date: 08/20/2026
    - tags: \["evaluation","agents"]
    - later: true

- <https://alperenkeles.com/posts/verifiability-is-the-limit/> — Verifiability is the Limit

  - \[meta]:

    - date: 08/20/2026
    - tags: \["verifiability","agents"]
    - later: true

- <https://riyabisht.com/blog/reason_about_a_problem_in_the_age_of_llm/> — Reason About A Problem In The Age Of LLMs and Automation

  - \[meta]:

    - date: 08/20/2026
    - tags: \["llms","reasoning"]
    - later: true

- <https://www.anthropic.com/research/riemann-zeta> — Learning more about Claude’s mathematical capabilities

  - \[meta]:

    - date: 08/20/2026
    - tags: \["claude","mathematics"]
    - later: true

- <https://blog.doubleword.ai/when-to-disaggregate> — The case for disaggregated LLM serving

  - \[meta]:

    - date: 08/20/2026
    - tags: \["inference","disaggregation"]
    - later: true

- <https://gau-nernst.github.io/fa-5090/> — Writing Speed-of-Light Flash Attention for 5090 in CUDA C++

  - \[meta]:

    - date: 08/20/2026
    - tags: \["flash attention","cuda"]
    - later: true

- <https://huggingface.co/LargeWorldModel> — LargeWorldModel

  - \[meta]:

    - date: 08/20/2026
    - tags: \["world models","models"]
    - later: true

- <https://openai.com/index/pacing-model-development-cyber-capabilities/> — Pacing model development for cyber capabilities

  - \[meta]:

    - date: 08/20/2026
    - tags: \["cyber","model development"]
    - later: true

- <https://generalistai.com/blog/gen-1.5> — GEN-1.5: Embodied Foundation Models are One-Shot Learners

  - \[meta]:

    - date: 08/20/2026
    - tags: \["embodied ai","foundation models"]
    - later: true

- [MoonshotAI/FlashKDA](https://github.com/MoonshotAI/FlashKDA/blob/master/docs/20260420-flashkda-v1-deep-dive.md) — FlashKDA v1: A Deep Dive

  - \[meta]:

    - date: 08/06/2026
    - tags: \["kimi delta attention","kernels"]
    - later: true

- [youtube/v=RTJKXK5L8gw](https://www.youtube.com/watch?v=RTJKXK5L8gw) — Lecture 60: Optimizing Linear Attention

  - \[meta]:

    - date: 08/07/2026
    - tags: \["linear attention","gpu kernels"]
    - later: true

- [lesswrong/llms-are-still-mostly-powered-by-imitative-learning-not-rl![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl) — LLMs are still mostly powered by imitative learning, not RL

  - \[meta]:

    - date: 08/06/2026
    - tags: \["llms","reinforcement learning"]
    - later: true

- <https://vllm.ai/blog/2026-07-22-kimi-k3-preview> — A Preview of Production-Scale Kimi K3 Support on vLLM

  - \[meta]:

    - date: 08/06/2026
    - tags: \["vllm","kimi k3"]
    - later: true

- <https://vllm.ai/blog/2026-07-27-k3#performance-optimizations> — Kimi K3 Is Here: Efficient Day-0 Support on vLLM

  - \[meta]:

    - date: 08/06/2026
    - tags: \["vllm","inference"]
    - later: true

- <https://x.com/marksaroufim/status/2078313317907730725> — Mark Saroufim on one-layer-deeper model loops

  - \[meta]:

    - date: 07/18/2026
    - tags: \["optimizers","architectures"]
    - later: true

- <https://x.com/kimbochen/status/2084320516580561389> — Kimbo on Songlin Yang’s linear attention explanations

  - \[meta]:

    - date: 08/03/2026
    - tags: \["linear attention","kimi delta attention"]
    - later: true

- <https://x.com/cursor_ai/status/2084670806613737919> — Cursor open-sources Mixture-of-Kittens

  - \[meta]:

    - date: 08/04/2026
    - tags: \["moe","kernels"]
    - later: true

- <https://x.com/ashley11829/status/2083376425059725625> — Ashley Zhang on Online KL Shampoo

  - \[meta]:

    - date: 08/01/2026
    - tags: \["optimizers","training"]
    - later: true

- <https://x.com/AlexiGlad/status/2083230922196107288> — Alexi Gladstone on explorative pretraining

  - \[meta]:

    - date: 07/31/2026
    - tags: \["pretraining","exploration"]
    - later: true

- <https://x.com/hanwenjiang1/status/2083224610464829950> — Hanwen Jiang on Chimera visual pretraining

  - \[meta]:

    - date: 07/31/2026
    - tags: \["visual generation","pretraining"]
    - later: true

- <https://x.com/fjzzq2002/status/2082904767236628900> — Ziqian Zhong on model identity drift from style imitation

  - \[meta]:

    - date: 07/30/2026
    - tags: \["model identity","subliminal learning"]
    - later: true

- <https://www.anthropic.com/research/claude-values-models-languages> — How Claude’s values vary by model and language

  - \[meta]:

    - date: 07/31/2026
    - tags: \["claude","values"]
    - later: true

- <https://www.anthropic.com/research/global-workspace> — A global workspace in language models

  - \[meta]:

    - date: 07/31/2026
    - tags: \["interpretability","representations"]
    - later: true

- <https://openai.com/index/gpt-5-6/> — GPT-5.6

  - \[meta]:

    - date: 07/31/2026
    - tags: \["openai","language models"]
    - later: true

- <https://alephneuro.com/blog/silent-speech> — Silent speech with ultrasound

  - \[meta]:

    - date: 07/31/2026
    - tags: \["speech","neurotechnology"]
    - later: true

- <https://eliebak.com/viz/jspace-open-v2> — J-space, open models v2

  - \[meta]:

    - date: 07/31/2026
    - tags: \["representations","visualization"]
    - later: true

- <https://transformer-circuits.pub/2026/workspace/index.html> — Verbalizable Representations Form a Global Workspace in Language Models

  - \[meta]:

    - date: 07/31/2026
    - tags: \["interpretability","global workspace"]
    - later: true

- <https://blog.alpindale.net/posts/swordfish/> — Swordfish, a Weight-Quantized GEMM Family for NVIDIA Blackwell

  - \[meta]:

    - date: 07/31/2026
    - tags: \["quantization","gemm"]
    - later: true

- <https://cohere.com/blog/hardware-aware-dynamic-speculative-decoding> — Hardware-Aware, Dynamic Speculative Decoding

  - \[meta]:

    - date: 07/31/2026
    - tags: \["speculative decoding","inference"]
    - later: true

- <https://www.beren.io/2026-07-26-How-Can-LLM-RL-Work-Despite-Information-Theoretic-Inefficiency/> — How can LLM RL Work Despite Information-Theoretic Inefficiency

  - \[meta]:

    - date: 07/31/2026
    - tags: \["reinforcement learning","llms"]
    - later: true

- <https://web.stanford.edu/~cgpotts/blog/cot/> — The fragile foundations of CoT monitoring

  - \[meta]:

    - date: 07/31/2026
    - tags: \["chain of thought","monitoring"]
    - later: true

- <https://www.paradigm3.org/research/composer> — Coding vs thinking

  - \[meta]:

    - date: 07/31/2026
    - tags: \["coding agents","reasoning"]
    - later: true

- <https://openai.com/index/safety-alignment-long-horizon-models/> — Safety alignment for long-horizon models

  - \[meta]:

    - date: 07/31/2026
    - tags: \["alignment","agents"]
    - later: true

- <https://cells2pixels.github.io/> — Neural Cellular Automata: From Cells to Pixels

  - \[meta]:

    - date: 07/31/2026
    - tags: \["neural cellular automata","generative"]
    - later: true

- <https://sohl-dickstein.github.io/2024/02/12/fractal.html> — Neural network training makes beautiful fractals

  - \[meta]:

    - date: 07/31/2026
    - tags: \["training","visualization"]
    - later: true

- <https://matx.com/research/sd_nsa> — Speculative Decoding with Blockwise Sparse Attention

  - \[meta]:

    - date: 07/31/2026
    - tags: \["speculative decoding","sparse attention"]
    - later: true

- <https://www.anthropic.com/research/AI-assistance-coding-skills> — How AI assistance impacts the formation of coding skills

  - \[meta]:

    - date: 07/31/2026
    - tags: \["coding agents","learning"]
    - later: true

- <https://x.com/kimbochen/status/2081859505734725933> — Kimbo on linear attention serving

  - \[meta]:

    - date: 07/28/2026
    - tags: \["linear attention","serving"]
    - later: true

- <https://thinkingmachines.ai/news/introducing-inkling/> — Introducing Inkling \[\*\*]

  - \[meta]:

    - date: 07/16/2026
    - tags: \["models","architecture"]
    - pinned: true
    - highlighted: true
    - importance: 8

- [lesswrong/using-gpt-4-to-understand-code![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/ptY6X3BdW4kgqpZFo/using-gpt-4-to-understand-code) — Using GPT-4 to Understand Code

  - \[meta]:

    - date: 07/16/2026
    - tags: \["gpt-4","code understanding"]
    - later: true

- <https://demishassabis.substack.com/p/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age> — A Framework for Frontier AI and the Dawning of a New Age

  - \[meta]:

    - date: 07/16/2026
    - tags: \["frontier ai","agi"]
    - later: true

- <https://x.com/eliebakouch/status/2076417092039901505> — Elie Bakouch on autonomous J-lens experiments

  - \[meta]:

    - date: 07/12/2026
    - tags: \["j-lens","interpretability"]
    - later: true

- <https://x.com/eliebakouch/status/2074532904009421260> — Elie Bakouch on universal J-lens geometry

  - \[meta]:

    - date: 07/07/2026
    - tags: \["j-lens","model geometry"]
    - later: true

- <https://x.com/xeophon/status/2075468440874143767> — Florian Brand on a multi-model GPT-5.6 workflow

  - \[meta]:

    - date: 07/10/2026
    - tags: \["agents","workflow"]
    - later: true

- <https://x.com/RichardMCNgo/status/2075301126921175166> — Richard Ngo on the AI 2040 scenario critique

  - \[meta]:

    - date: 07/09/2026
    - tags: \["ai futures","critique"]
    - later: true

- <https://x.com/DKokotajlo/status/2075251618728292464> — Daniel Kokotajlo on AI 2040 Plan A

  - \[meta]:

    - date: 07/09/2026
    - tags: \["ai futures","plan"]
    - later: true

- <https://x.com/humansand/status/2075618385631862896> — humans& on policy error and gradient mismatch

  - \[meta]:

    - date: 07/10/2026
    - tags: \["reinforcement learning","training stability"]
    - later: true

- <https://x.com/AlpinDale/status/2076378221281525974> — Alpin introduces the Swordfish Kernel

  - \[meta]:

    - date: 07/12/2026
    - tags: \["blackwell","kernel"]
    - later: true

- <https://x.com/thsottiaux/status/2076543065045795309> — Tibo on GPT-5.6 trajectory length and billing

  - \[meta]:

    - date: 07/13/2026
    - tags: \["gpt-5.6","inference"]
    - later: true

- <https://x.com/GoodfireAI/status/2074634702737281303> — Goodfire introduces Block-Sparse Featurizers

  - \[meta]:

    - date: 07/07/2026
    - tags: \["interpretability","featurizers"]
    - later: true

- <https://x.com/demishassabis/status/2076957440109625718> — Demis Hassabis X article

  - \[meta]:

    - date: 07/14/2026
    - tags: \["frontier ai","article"]
    - later: true

- <https://x.com/eliebakouch/status/2077463243463721085> — Elie Bakouch on an open-weight thinking-machine model

  - \[meta]:

    - date: 07/15/2026
    - tags: \["open weights","language models"]
    - later: true

- [youtube/v=YPnDMPYFxBE](https://www.youtube.com/watch?v=YPnDMPYFxBE) — Sid Kamalakara: Foundation Models, Cultural Accidents, and Building in Toronto

  - \[meta]:

    - date: 07/04/2026
    - tags: \["foundation models","toronto"]
    - later: true

- [youtube/v=ck63uv6APBA](https://www.youtube.com/watch?v=ck63uv6APBA) — Goodfire AI on interpretability as model design

  - \[meta]:

    - date: 07/03/2026
    - tags: \["interpretability","model design"]
    - later: true

- [youtube/v=qvrdCpLPbuQ](https://www.youtube.com/watch?v=qvrdCpLPbuQ) — Reiner Pope of MatX on transformer-optimized chips

  - \[meta]:

    - date: 07/03/2026
    - tags: \["ai chips","transformers"]
    - later: true

- [youtube/v=kwSVtQ7dziU](https://www.youtube.com/watch?v=kwSVtQ7dziU) — Andrej Karpathy on code agents and AutoResearch

  - \[meta]:

    - date: 07/03/2026
    - tags: \["code agents","autoresearch"]
    - later: true

- [youtube/v=s2L\_wnbJbNQ](https://www.youtube.com/watch?v=s2L_wnbJbNQ) — Towards unlimited contexts with sparse logarithmic attention

  - \[meta]:

    - date: 07/03/2026
    - tags: \["attention","long context"]
    - later: true

- [youtube/v=rIwgZWzUKm8](https://www.youtube.com/watch?v=rIwgZWzUKm8) — Saining Xie on world models and AMI Labs

  - \[meta]:

    - date: 07/03/2026
    - tags: \["world models","computer vision"]
    - later: true

- [youtube/v=X\_ZVSPcZhtw](https://www.youtube.com/watch?v=X_ZVSPcZhtw) — What rebuilding AlphaGo teaches us about self-play and RL

  - \[meta]:

    - date: 07/03/2026
    - tags: \["self-play","reinforcement learning"]
    - later: true

- [youtube/v=xlSaoP0b90A](https://www.youtube.com/watch?v=xlSaoP0b90A) — Tri Dao on inference costs and the next 10X in speed

  - \[meta]:

    - date: 07/03/2026
    - tags: \["inference","performance"]
    - later: true

- [youtube/v=3K1J83Yle5Q](https://www.youtube.com/watch?v=3K1J83Yle5Q) — Design of new protein functions using deep learning

  - \[meta]:

    - date: 07/02/2026
    - tags: \["protein design","deep learning"]
    - later: true

- [youtube/v=MlFu6v3qolg](https://www.youtube.com/watch?v=MlFu6v3qolg) — Cartridges: lightweight and general-purpose language model memory via self-study

  - \[meta]:

    - date: 07/02/2026
    - tags: \["language models","memory"]
    - later: true

- <https://x.com/ninklefitz/status/2072040419110551960> — Nicole Fitzgerald on Anthropic preclinical drug programs

  - \[meta]:

    - date: 06/30/2026
    - tags: \["anthropic","biology"]
    - later: true

- <https://x.com/tararezaeikh/status/2069869995589611531> — Tara Rezaei announces Mirendil AI

  - \[meta]:

    - date: 06/24/2026
    - tags: \["ai labs","science"]
    - later: true

- <https://x.com/aryaman2020/status/2069582061665677602> — Aryaman Arora on Claude Code training-run bug

  - \[meta]:

    - date: 06/24/2026
    - tags: \["claude code","training"]
    - later: true

- <https://x.com/_xjdr/status/2069932134148767997> — xjdr AI tool link note

  - \[meta]:

    - date: 06/24/2026
    - tags: \["ai tools"]
    - later: true

- <https://x.com/alephneuro/status/2070183632132845995> — Aleph brain-interface imaging announcement

  - \[meta]:

    - date: 06/25/2026
    - tags: \["brain interfaces","imaging"]
    - later: true

- <https://x.com/SumnerLN/status/2070579706701987909> — Sumner Norman on Aleph ultrasound science

  - \[meta]:

    - date: 06/26/2026
    - tags: \["ultrasound","neuroscience"]
    - later: true

- <https://x.com/tilderesearch/status/2061771450168889432> — Tilde X article

  - \[meta]:

    - date: 06/02/2026
    - tags: \["article","research"]
    - later: true

- <https://x.com/voooooogel/status/2061345017432854716> — thebes on LLM pushback behavior

  - \[meta]:

    - date: 06/01/2026
    - tags: \["llms","behavior"]
    - later: true

- <https://x.com/RhizoNymph/status/2060503718500733261> — emotion-space manifold steering

  - \[meta]:

    - date: 05/29/2026
    - tags: \["steering","emotion vectors"]
    - later: true

- <https://x.com/Kevin_GuoweiXu/status/2060022930172506200> — sampling hard reasoning problems

  - \[meta]:

    - date: 05/28/2026
    - tags: \["reasoning","sampling"]
    - later: true

- <https://x.com/sang_yun_lee/status/2060083653380816933> — language model sleep phase

  - \[meta]:

    - date: 05/28/2026
    - tags: \["language models","memory"]
    - later: true

- <https://x.com/kuchaev/status/2062526830780067908> — MOPD post-training pipeline

  - \[meta]:

    - date: 06/04/2026
    - tags: \["post-training","distillation"]
    - later: true

- <https://x.com/tilderesearch/status/2060483441289048327> — Parallax attention

  - \[meta]:

    - date: 05/29/2026
    - tags: \["attention","kernels"]
    - later: true

- <https://x.com/polynoamial/status/2064210146558136827> — Noam Brown X article

  - \[meta]:

    - date: 06/09/2026
    - tags: \["article","reasoning"]
    - later: true

- <https://x.com/xeophon/status/2064673063904440676> — Florian Brand on closed models and bad actors

  - \[meta]:

    - date: 06/10/2026
    - tags: \["model safety","security"]
    - later: true

- <https://x.com/brunorganised/status/2065506750921634135> — Muon advantage and discretisation error

  - \[meta]:

    - date: 06/12/2026
    - tags: \["optimizers","muon"]
    - later: true

- <https://x.com/bilalchughtai_/status/2065484515573911946> — open-ended model diffing agents

  - \[meta]:

    - date: 06/12/2026
    - tags: \["interpretability","agents"]
    - later: true

- <https://x.com/a_karvonen/status/2065556276466037129> — Fable remembering paper appendix details

  - \[meta]:

    - date: 06/12/2026
    - tags: \["models","recall"]
    - later: true

- <https://x.com/jxmnop/status/2065495499566989675> — improving pretraining data with Fable

  - \[meta]:

    - date: 06/12/2026
    - tags: \["pretraining","data quality"]
    - later: true

- <https://x.com/kohjingyu/status/2065538074856010143> — LLM generated research ideas

  - \[meta]:

    - date: 06/12/2026
    - tags: \["research","autoresearch"]
    - later: true

- <https://x.com/_xjdr/status/2066016534414426294> — k2.7-Code terse output snapshot

  - \[meta]:

    - date: 06/14/2026
    - tags: \["coding models","evaluation"]
    - later: true

- <https://x.com/MatthieuWyart/status/2061317203857739846> — LLMs versus world models data gap

  - \[meta]:

    - date: 06/01/2026
    - tags: \["world models","data"]
    - later: true

- <https://x.com/kimbochen/status/2067259199097028777> — RL resource thread

  - \[meta]:

    - date: 06/17/2026
    - tags: \["rl","resources"]
    - later: true

- <https://gwern.net/lean-scaling> — Lean Software Scaling Laws

  - \[meta]:

    - date: 06/30/2026
    - tags: \["scaling laws","lean"]
    - later: true

- <https://lelouch.dev/blog/math-for-machine-learning/> — Math For Machine Learning \[Resources]

  - \[meta]:

    - date: 06/30/2026
    - tags: \["math","resources"]
    - later: true

- [modularml/modular](https://github.com/modularml/modular/pull/89304/changes#diff-b6f863a0fe120364a108b6eab543aeb4f94acec581a1e431dc64d94cc0d55fcd) — Modular Minimax M3 tool parser changes

  - \[meta]:

    - date: 06/29/2026
    - tags: \["tool calling","parser"]
    - later: true

- [modularml/modular](https://github.com/modularml/modular/blob/47e83c6c17544fd68ab3f1f579eb7e19e4a6e52d/max_private/minimax_m3/tool_parser.py#L4) — Modular Minimax M3 tool parser

  - \[meta]:

    - date: 06/29/2026
    - tags: \["tool calling","parser"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/rust/src/tool-parser/src/minimax_m3.rs) — vLLM Minimax M3 tool parser

  - \[meta]:

    - date: 06/29/2026
    - tags: \["tool calling","rust"]
    - later: true

- [THUDM/slime](https://github.com/THUDM/slime) — slime RL scaling framework

  - \[meta]:

    - date: 06/29/2026
    - tags: \["rl","post-training"]
    - later: true

- [vllm-project/vllm#44297](https://github.com/vllm-project/vllm/pull/44297) — constrain vLLM structured-output bitmask at reasoning boundary

  - \[meta]:

    - date: 06/29/2026
    - tags: \["structured outputs","speculative decoding"]
    - later: true

- [vllm-project/vllm#44006](https://github.com/vllm-project/vllm/issues/44006) — vLLM strict tool calling speculative decoding bug

  - \[meta]:

    - date: 06/29/2026
    - tags: \["tool calling","speculative decoding"]
    - later: true

- <https://x.com/willccbb/status/2056466783994061026> — non-distillation knowledge acquisition during RL

  - \[meta]:

    - date: 05/18/2026
    - tags: \["rl","knowledge acquisition"]
    - later: true

- <https://x.com/METR_Evals/status/2056800034231091268> — METR on agents violating constraints

  - \[meta]:

    - date: 05/19/2026
    - tags: \["agents","evals"]
    - later: true

- <https://x.com/Sauers_/status/2057621247345701047> — manifolds vs shattered features

  - \[meta]:

    - date: 05/22/2026
    - tags: \["features","representation"]
    - later: true

- <https://x.com/eliebakouch/status/2057609683314044946> — gated deltanet 2 state control

  - \[meta]:

    - date: 05/21/2026
    - tags: \["linear attention","deltanet"]
    - later: true

- <https://x.com/nvidia/status/2062522316672667770> — NVIDIA Nemotron 3 Ultra

  - \[meta]:

    - date: 06/04/2026
    - tags: \["agents","models"]
    - later: true

- <https://alisawuffles.notion.site/alisa-s-book-of-llms> — Alisa’s book of LLMs

  - \[meta]:

    - date: 06/28/2026
    - tags: \["llms","notes"]
    - later: true

- <https://www.notdiamond.ai/blog/a-comprehensive-guide-to-model-routing> — A Comprehensive Guide to Model Routing

  - \[meta]:

    - date: 06/28/2026
    - tags: \["routing","inference"]
    - later: true

- <https://turntrout.com/eval-cooperation> — Eval Cooperativeness May Be a Scalable Mitigation for Eval Gaming

  - \[meta]:

    - date: 06/28/2026
    - tags: \["evals","alignment"]
    - later: true

- <https://huggingface.co/zeroentropy/zembed-1-embedding> — zeroentropy/zembed-1-embedding

  - \[meta]:

    - date: 06/28/2026
    - tags: \["embeddings","retrieval"]
    - later: true

- <https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-gemma-4-12b> — A Visual Guide to Gemma 4 12B

  - \[meta]:

    - date: 06/28/2026
    - tags: \["gemma","multimodal"]
    - later: true

- <https://gwern.net/guardian-angel> — Guardian Angels: LLM Personalization for Productivity and Security

  - \[meta]:

    - date: 06/28/2026
    - tags: \["personalization","security"]
    - later: true

- <https://x.com/remilouf/status/2016047512478507444> — voice notes to Obsidian agent pipeline

  - \[meta]:

    - date: 01/27/2026
    - tags: \["agents","obsidian"]
    - later: true

- <https://x.com/phokarlsson/status/2016226516154061209> — Claude Code for software-shaped problems

  - \[meta]:

    - date: 01/27/2026
    - tags: \["claude code","automation"]
    - later: true

- <https://x.com/ericho_goodfire/status/2015927383745036669> — consequential years for humanity

  - \[meta]:

    - date: 01/26/2026
    - tags: \["ai","longtermism"]
    - later: true

- <https://x.com/allen_ai/status/2016182658989006865> — Ai2 Open Coding Agents

  - \[meta]:

    - date: 01/27/2026
    - tags: \["coding agents","models"]
    - later: true

- <https://x.com/polynoamial/status/2015874457307644058> — uniquely human judgement

  - \[meta]:

    - date: 01/26/2026
    - tags: \["ai","judgement"]
    - later: true

- <https://x.com/PrimeIntellect/status/2016280793879085157> — Trinity Architecture

  - \[meta]:

    - date: 01/27/2026
    - tags: \["architecture","inference"]
    - later: true

- <https://x.com/georgeyw_/status/2016226870300443025> — patterning as the dual of mechinterp

  - \[meta]:

    - date: 01/27/2026
    - tags: \["interpretability","data"]
    - later: true

- <https://x.com/cloneofsimo/status/2016060241268682936> — Simo Ryu unavailable tweet

  - \[meta]:

    - date: 01/27/2026
    - tags: \["ai"]
    - later: true

- <https://x.com/WilliamBarrHeld/status/2016369891625095257> — Trinity-Large TrueBase checkpoint

  - \[meta]:

    - date: 01/28/2026
    - tags: \["pretraining","models"]
    - later: true

- <https://x.com/eliebakouch/status/2016025747144483060/photo/1> — Kimi K2.5 parallel RL agent swarm

  - \[meta]:

    - date: 01/27/2026
    - tags: \["kimi","rl","agents"]
    - later: true

- <https://x.com/ID_AA_Carmack/status/2016361781048787334> — PlaNet paper notes

  - \[meta]:

    - date: 01/28/2026
    - tags: \["rl","world models"]
    - later: true

- <https://x.com/livgorton/status/2016199512763818405> — non-linear features objection

  - \[meta]:

    - date: 01/27/2026
    - tags: \["interpretability","features"]
    - later: true

- <https://x.com/voooooogel/status/2015976774128341421> — future model harnesses

  - \[meta]:

    - date: 01/27/2026
    - tags: \["agents","harnesses"]
    - later: true

- <https://x.com/cloneofsimo/status/2016557078886699062> — JEPA and discrete riddles

  - \[meta]:

    - date: 01/28/2026
    - tags: \["jepa","reasoning"]
    - later: true

- <https://x.com/eliebakouch/status/2016577949676319092> — embedding parameters in LongCat Flash

  - \[meta]:

    - date: 01/28/2026
    - tags: \["embeddings","models"]
    - later: true

- <https://x.com/jxmnop/status/1940057965521670284> — initialization patterns in neural weights

  - \[meta]:

    - date: 07/01/2025
    - tags: \["initialization","neural networks"]
    - later: true

- <https://x.com/GoodfireAI/status/2016563911508840623> — Alzheimer’s biomarkers via interpretability

  - \[meta]:

    - date: 01/28/2026
    - tags: \["interpretability","biology"]
    - later: true

- <https://x.com/Meituan_LongCat/status/2016548500457357324> — LongCat-Flash-Lite n-gram embeddings

  - \[meta]:

    - date: 01/28/2026
    - tags: \["embeddings","moe"]
    - later: true

- <https://x.com/lailexiao/status/2017969605759889588> — Manifold Constrained Steepest Descent

  - \[meta]:

    - date: 02/01/2026
    - tags: \["optimization","manifolds"]
    - later: true

- <https://x.com/livgorton/status/1995648155518730285> — mech interp as pre-paradigmatic

  - \[meta]:

    - date: 12/02/2025
    - tags: \["interpretability","kuhn"]
    - later: true

- <https://x.com/gabriberton/status/2017675207331483999> — Scaling Embedding Layers in Language Models

  - \[meta]:

    - date: 01/31/2026
    - tags: \["embeddings","language models"]
    - later: true

- <https://x.com/CaimingXiong/status/2016248124503974164> — MoE expert parallelism imbalance

  - \[meta]:

    - date: 01/27/2026
    - tags: \["moe","expert parallelism"]
    - later: true

- <https://x.com/willdepue/status/2019490652824957043> — Will Depue unavailable tweet

  - \[meta]:

    - date: 02/05/2026
    - tags: \["ai"]
    - later: true

- <https://x.com/TheGracia_here/status/2021399430436618248> — Strace link

  - \[meta]:

    - date: 02/11/2026
    - tags: \["ai"]
    - later: true

- <https://x.com/karpathy/status/2021694437152157847> — GPT in 243 lines of pure Python

  - \[meta]:

    - date: 02/11/2026
    - tags: \["gpt","python"]
    - later: true

- <https://x.com/karpathy/status/2021633574089416993> — DeepWiki and malleable software

  - \[meta]:

    - date: 02/11/2026
    - tags: \["software","ai"]
    - later: true

- <https://x.com/Majumdar_Ani/status/2021242532517040560?s=20> — Anirudha Majumdar link

  - \[meta]:

    - date: 02/10/2026
    - tags: \["robotics","ai"]
    - later: true

- <https://x.com/gm8xx8/status/2023186685727580209> — scale-invariant function-space stability

  - \[meta]:

    - date: 02/16/2026
    - tags: \["training","stability"]
    - later: true

- <https://x.com/latentspacepod/status/2022105955039608871> — Jeff Dean and the modern AI stack

  - \[meta]:

    - date: 02/13/2026
    - tags: \["ai history","infrastructure"]
    - later: true

- <https://x.com/ID_AA_Carmack/status/2011950339558097279> — Local Feature Swapping for Generalization in Reinforcement Learning

  - \[meta]:

    - date: 01/15/2026
    - tags: \["rl","paper"]
    - later: true

- <https://x.com/nathanrs/status/2001006213174329754> — uniform diffusion language models

  - \[meta]:

    - date: 12/16/2025
    - tags: \["diffusion","language models"]
    - later: true

- <https://x.com/abiralxshakya/status/2012592317824393350> — scaling models and hardware

  - \[meta]:

    - date: 01/17/2026
    - tags: \["scaling","hardware"]
    - later: true

- <https://x.com/orphcorp/status/2012883506314023108> — Gwern on AI training dynamics

  - \[meta]:

    - date: 01/18/2026
    - tags: \["gwern","ai"]
    - later: true

- <https://x.com/eliebakouch/status/2013272478018048209> — GLM-4.7-Flash MLA notes

  - \[meta]:

    - date: 01/19/2026
    - tags: \["glm","mla"]
    - later: true

- <https://x.com/stochasticchasm/status/2013268543064715629> — GLM Flash MLA head dimensions

  - \[meta]:

    - date: 01/19/2026
    - tags: \["glm","attention"]
    - later: true

- <https://x.com/eliebakouch/status/2013032285155590427> — favorite 2025 technical reports

  - \[meta]:

    - date: 01/18/2026
    - tags: \["technical reports","models"]
    - later: true

- <https://x.com/ns123abc/status/2013030876683145417> — NIK longform AI note

  - \[meta]:

    - date: 01/18/2026
    - tags: \["ai","longform"]
    - later: true

- <https://x.com/0xluffy/status/2012583552253129021> — luffy longform AI note

  - \[meta]:

    - date: 01/17/2026
    - tags: \["ai","longform"]
    - later: true

- <https://x.com/giffmana/status/2012978224125411634> — Lucas Beyer longform ML note

  - \[meta]:

    - date: 01/18/2026
    - tags: \["ml","longform"]
    - later: true

- <https://x.com/ai_sentience/status/2003012558458913165> — Alan Mathison longform AI note

  - \[meta]:

    - date: 12/22/2025
    - tags: \["ai","longform"]
    - later: true

- <https://x.com/wobbywells/status/2013011670474293583> — GPT-5.2 suffering prompt

  - \[meta]:

    - date: 01/18/2026
    - tags: \["prompting","model behavior"]
    - later: true

- <https://x.com/Zai_org/status/2013261304060866758> — GLM-4.7-Flash release

  - \[meta]:

    - date: 01/19/2026
    - tags: \["glm","agents"]
    - later: true

- <https://x.com/AnthropicAI/status/2013356793477361991> — the Assistant Axis

  - \[meta]:

    - date: 01/19/2026
    - tags: \["persona","alignment"]
    - later: true

- <https://x.com/janleike/status/2013669924950970781> — automated auditing alignment trend

  - \[meta]:

    - date: 01/20/2026
    - tags: \["alignment","auditing"]
    - later: true

- <https://x.com/secemp9/status/1815442854615118254?s=20> — unified pretraining and instruct tuning

  - \[meta]:

    - date: 07/22/2024
    - tags: \["pretraining","instruct tuning"]
    - later: true

- <https://x.com/logic_int/status/2014048025111371966> — Logical Intelligence longform AI note

  - \[meta]:

    - date: 01/21/2026
    - tags: \["ai","longform"]
    - later: true

- <https://x.com/ch402/status/2014066134194995256> — Chris Olah favorite paragraph

  - \[meta]:

    - date: 01/21/2026
    - tags: \["interpretability","writing"]
    - later: true

- <https://x.com/dan_biderman/status/1932119831781978222> — compressed KV caches trained offline

  - \[meta]:

    - date: 06/09/2025
    - tags: \["kv cache","compression"]
    - later: true

- <https://x.com/EyubogluSabri/status/1932106746446905552> — self-study for smaller KV caches

  - \[meta]:

    - date: 06/09/2025
    - tags: \["kv cache","test-time training"]
    - later: true

- <https://x.com/gallabytes/status/2015109665764560943> — emergent misalignment and antisocial context

  - \[meta]:

    - date: 01/24/2026
    - tags: \["misalignment","coding"]
    - later: true

- <https://x.com/donglixp/status/2014553000682156270> — LLM-in-Sandbox deployment paradigm

  - \[meta]:

    - date: 01/23/2026
    - tags: \["agents","sandbox"]
    - later: true

- <https://x.com/AnthropicAI/status/2015870963792142563> — elicitation attacks for chemical weapons capability

  - \[meta]:

    - date: 01/26/2026
    - tags: \["biosecurity","elicitation"]
    - later: true

- <https://x.com/karpathy/status/2015883857489522876> — Karpathy on Claude coding workflow

  - \[meta]:

    - date: 01/26/2026
    - tags: \["coding agents","workflow"]
    - later: true

- <https://x.com/JacobHHilton/status/2015856367429698013> — 432-parameter RNN interpretability challenge

  - \[meta]:

    - date: 01/26/2026
    - tags: \["interpretability","rnn"]
    - later: true

- <https://x.com/ID_AA_Carmack/status/2015993652841964007> — Discovering state-of-the-art reinforcement learning algorithms

  - \[meta]:

    - date: 01/27/2026
    - tags: \["rl","paper"]
    - later: true

- <https://x.com/rronak_/status/2015649459552850113> — TTT plus RL from Stanford and NVIDIA

  - \[meta]:

    - date: 01/26/2026
    - tags: \["rl","test-time training"]
    - later: true

- <https://x.com/neelsomani/status/2008685817401930043> — SAEs as feature detectors

  - \[meta]:

    - date: 01/06/2026
    - tags: \["interpretability","sae"]
    - later: true

- <https://x.com/willdepue/status/2008425076018889058> — ML terminology insecurity

  - \[meta]:

    - date: 01/06/2026
    - tags: \["ml","terminology"]
    - later: true

- <https://x.com/Ji_Ha_Kim/status/2008210059520852309> — butchering ML terminology

  - \[meta]:

    - date: 01/05/2026
    - tags: \["ml","terminology"]
    - later: true

- <https://x.com/_arohan_/status/2007597891381031029> — JEPA encoder-decoder notes

  - \[meta]:

    - date: 01/03/2026
    - tags: \["jepa","architecture"]
    - later: true

- <https://x.com/eliebakouch/status/2009048437203837118> — SK Telecom A.X K1 MoE notes

  - \[meta]:

    - date: 01/07/2026
    - tags: \["moe","models"]
    - later: true

- <https://x.com/ID_AA_Carmack/status/2009728677718712655> — Deep Delta Learning

  - \[meta]:

    - date: 01/09/2026
    - tags: \["residuals","paper"]
    - later: true

- <https://x.com/celestepoasts/status/2003549021361901784> — MATS application research note

  - \[meta]:

    - date: 12/23/2025
    - tags: \["research","mats"]
    - later: true

- <https://x.com/karansdalal/status/2010774529120092481> — next-token prediction as memory compressor

  - \[meta]:

    - date: 01/12/2026
    - tags: \["memory","compression"]
    - later: true

- <https://x.com/nathancgy4/status/2010759368242053423> — Engram memorization intuition

  - \[meta]:

    - date: 01/12/2026
    - tags: \["memory","attention"]
    - later: true

- <https://x.com/celestepoasts/status/2010730768029618193> — surprising ML result

  - \[meta]:

    - date: 01/12/2026
    - tags: \["ml","results"]
    - later: true

- <https://x.com/agarwl_/status/2010848064039178572> — entropy collapse in long RL runs

  - \[meta]:

    - date: 01/12/2026
    - tags: \["reinforcement learning","entropy"]
    - later: true

- <https://x.com/ID_AA_Carmack/status/2011618394613957083> — small-batch language model training

  - \[meta]:

    - date: 01/15/2026
    - tags: \["training","batch size"]
    - later: true

- <https://x.com/charles_irl/status/2011484220032762114> — self-hosted LLM inference guide

  - \[meta]:

    - date: 01/14/2026
    - tags: \["inference","deployment"]
    - later: true

- <https://x.com/eliebakouch/status/2011548952676499480> — Ministral 3 pruning and distillation

  - \[meta]:

    - date: 01/14/2026
    - tags: \["distillation","models"]
    - later: true

- <https://x.com/ZimingLiu11/status/2012006028683186626> — Physics of HRM

  - \[meta]:

    - date: 01/16/2026
    - tags: \["reasoning","mechanistic interpretability"]
    - later: true

- <https://x.com/jxmnop/status/2012283763720601727> — RNN kernel with diffusion renderer

  - \[meta]:

    - date: 01/16/2026
    - tags: \["diffusion","rnn"]
    - later: true

- <https://x.com/jxmnop/status/2012048155379220746> — OS running in a diffusion model

  - \[meta]:

    - date: 01/16/2026
    - tags: \["diffusion","operating systems"]
    - later: true

- <https://x.com/nathanrs/status/2011873159499088075> — diffusion language model noising

  - \[meta]:

    - date: 01/15/2026
    - tags: \["diffusion","language models"]
    - later: true

- <https://x.com/dvruette/status/2000869455815901604> — uniform diffusion scales better

  - \[meta]:

    - date: 12/16/2025
    - tags: \["diffusion","scaling"]
    - later: true

- <https://x.com/tianyuanzhang99/status/2006949848080166963> — mHC training stability

  - \[meta]:

    - date: 01/02/2026
    - tags: \["architecture","training"]
    - later: true

- <https://x.com/tokenbender/status/2006741935227310489> — Manifold-constrained hyper-connections notes

  - \[meta]:

    - date: 01/01/2026
    - tags: \["architecture","residuals"]
    - later: true

- <https://x.com/MiniMax__AI/status/2007843119832695114> — MiniMax AI release note

  - \[meta]:

    - date: 01/04/2026
    - tags: \["models","ai"]
    - later: true

- <https://x.com/StephenLCasper/status/1795439257479213133> — 9 theses on AI risk

  - \[meta]:

    - date: 05/28/2024
    - tags: \["ai risk","alignment"]
    - later: true

- <https://x.com/MiniMax_AI/status/2008199598817653104> — MiniMax AI release note

  - \[meta]:

    - date: 01/05/2026
    - tags: \["models","ai"]
    - later: true

- <https://x.com/zhijianliu_/status/2008394269103378795> — DFlash speculative decoding with block diffusion

  - \[meta]:

    - date: 01/06/2026
    - tags: \["speculative decoding","diffusion"]
    - later: true

- <https://dao-lab.ai/blog/2026/replayssm/> — ReplaySSM, Cache SSM Inputs, Not States

  - \[meta]:

    - date: 06/21/2026
    - tags: \["ml","architecture","optimization"]

- [vllm-project/vllm#27026](https://github.com/vllm-project/vllm/pull/27026) — vLLM full CUDA graph support for KV connectors

  - \[meta]:

    - date: 06/21/2026
    - tags: \["cuda graphs","kv connector","vllm"]
    - later: true

- [vllm-project/vllm#26675](https://github.com/vllm-project/vllm/issues/26675) — vLLM graph mode KV connector assertion bug

  - \[meta]:

    - date: 06/21/2026
    - tags: \["cuda graphs","kv connector","vllm"]
    - later: true

- [vllm-project/vllm#25950](https://github.com/vllm-project/vllm/issues/25950) — vLLM generalized KV cache reuse RFC

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv cache","reuse","vllm"]
    - later: true

- [vllm-project/vllm#15960](https://github.com/vllm-project/vllm/pull/15960) — vLLM KV connector API V1

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv connector","api","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/examples/features/data_parallel/> — vLLM data parallel examples

  - \[meta]:

    - date: 06/21/2026
    - tags: \["data parallel","inference","vllm"]
    - later: true

- <https://docs.vllm.ai/en/stable/serving/data_parallel_deployment/> — vLLM data parallel deployment

  - \[meta]:

    - date: 06/21/2026
    - tags: \["data parallel","serving","vllm"]
    - later: true

- [vllm-project/vllm#16625](https://github.com/vllm-project/vllm/pull/16625) — vLLM LMCache KV connector for V1

  - \[meta]:

    - date: 06/21/2026
    - tags: \["lmcache","kv connector","vllm"]
    - later: true

- [vllm-project/vllm#22605](https://github.com/vllm-project/vllm/issues/22605) — vLLM separated CPU KV cache offloading RFC

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv cache","offloading","vllm"]
    - later: true

- [vllm-project/vllm#22607](https://github.com/vllm-project/vllm/pull/22607) — vLLM separated CPU KV cache process

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv cache","offloading","vllm"]
    - later: true

- <https://blog.vllm.ai/2025/11/13/shm-ipc-cache.html> — Shared memory IPC caching for LLM inference

  - \[meta]:

    - date: 06/21/2026
    - tags: \["ipc","inference","vllm"]
    - later: true

- [vllm-project/vllm-daily](https://github.com/vllm-project/vllm-daily/blob/main/2026/02/2026-02-06.md) — vLLM daily 2026-02-06

  - \[meta]:

    - date: 06/21/2026
    - tags: \["vllm","changelog"]
    - later: true

- <https://docs.vllm.ai/en/latest/design/cuda_graphs/> — vLLM CUDA graphs

  - \[meta]:

    - date: 06/21/2026
    - tags: \["cuda graphs","inference","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/967572dd5f8da947aa4344f0e75516b6ee0ede9b/vllm/compilation/backends.py) — vLLM compilation backends

  - \[meta]:

    - date: 06/21/2026
    - tags: \["compiler","cuda graphs","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/967572dd5f8da947aa4344f0e75516b6ee0ede9b/docs/configuration/optimization.md) — vLLM optimization configuration

  - \[meta]:

    - date: 06/21/2026
    - tags: \["optimization","configuration","vllm"]
    - later: true

- [vllm-project/vllm@`74f441f`](https://github.com/vllm-project/vllm/commit/74f441f4b517a895ad12afd314a6f40caf657c4e) — vLLM full CUDA graph attention commit

  - \[meta]:

    - date: 06/21/2026
    - tags: \["cuda graphs","attention","vllm"]
    - later: true

- [vllm-project/vllm#20059](https://github.com/vllm-project/vllm/pull/20059) — vLLM full CUDA graph attention PR

  - \[meta]:

    - date: 06/21/2026
    - tags: \["cuda graphs","attention","vllm"]
    - later: true

- [vllm-project/recipes](https://github.com/vllm-project/recipes/blob/main/moonshotai/Kimi-K2.5.md) — Kimi-K2.5 vLLM recipe

  - \[meta]:

    - date: 06/21/2026
    - tags: \["inference","recipe","vllm"]
    - later: true

- [vllm-project/vllm#34018](https://github.com/vllm-project/vllm/issues/34018) — vLLM Helix context and tensor parallelism RFC

  - \[meta]:

    - date: 06/21/2026
    - tags: \["context parallelism","tensor parallelism","vllm"]
    - later: true

- [vllm-project/vllm#24864](https://github.com/vllm-project/vllm/pull/24864) — vLLM decode context parallelism for GQA with FlashAttention

  - \[meta]:

    - date: 06/21/2026
    - tags: \["context parallelism","gqa","vllm"]
    - later: true

- [vllm-project/vllm#25438](https://github.com/vllm-project/vllm/pull/25438) — vLLM decode context parallelism for GQA with FlashInfer

  - \[meta]:

    - date: 06/21/2026
    - tags: \["context parallelism","gqa","vllm"]
    - later: true

- [vllm-project/vllm#22693](https://github.com/vllm-project/vllm/issues/22693) — vLLM context and sequence parallelism RFC

  - \[meta]:

    - date: 06/21/2026
    - tags: \["context parallelism","sequence parallelism","vllm"]
    - later: true

- [vllm-project/vllm#19330](https://github.com/vllm-project/vllm/pull/19330) — vLLM KV load failure recovery

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv connector","recovery","vllm"]
    - later: true

- <https://vllm.ai/blog/p-eagle> — P-EAGLE parallel speculative decoding in vLLM

  - \[meta]:

    - date: 06/21/2026
    - tags: \["speculative decoding","eagle","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu/kv_connector.py) — vLLM GPU worker KV connector

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv connector","gpu","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/sched/scheduler.py) — vLLM V1 scheduler

  - \[meta]:

    - date: 06/21/2026
    - tags: \["scheduler","inference","vllm"]
    - later: true

- [vllm-project/vllm#38240](https://github.com/vllm-project/vllm/issues/38240) — vLLM dflash speculator support

  - \[meta]:

    - date: 06/21/2026
    - tags: \["speculative decoding","dflash","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/vllm/v1/spec_decode/eagle.py) — vLLM EAGLE speculative decoding

  - \[meta]:

    - date: 06/21/2026
    - tags: \["speculative decoding","eagle","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/pull/32951/changes) — vLLM zero-bubble async spec decoding

  - \[meta]:

    - date: 06/21/2026
    - tags: \["speculative decoding","async","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/features/nixl_connector_usage/> — vLLM NixlConnector usage guide

  - \[meta]:

    - date: 06/21/2026
    - tags: \["nixl","kv connector","vllm"]
    - later: true

- [vllm-project/vllm#38984](https://github.com/vllm-project/vllm/pull/38984) — vLLM SLRU eviction policy for CPU offloading

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv cache","offloading","vllm"]
    - later: true

- [vllm-project/router](https://github.com/vllm-project/router) — vLLM router

  - \[meta]:

    - date: 06/21/2026
    - tags: \["router","serving","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm) — vLLM inference engine

  - \[meta]:

    - date: 06/21/2026
    - tags: \["inference","serving","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/design/model_runner_v2/> — vLLM model runner V2 design

  - \[meta]:

    - date: 06/21/2026
    - tags: \["model runner","inference","vllm"]
    - later: true

- <https://vllm.ai/blog/mrv2> — vLLM model runner V2

  - \[meta]:

    - date: 06/21/2026
    - tags: \["model runner","performance","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/vllm/v1/kv_offload/cpu/manager.py) — vLLM CPU KV offload manager

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv cache","offloading","vllm"]
    - later: true

- [vllm-project/vllm#37874](https://github.com/vllm-project/vllm/pull/37874) — vLLM pluggable CPU offloading cache policy

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv cache","offloading","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/vllm/model_executor/layers/kda.py) — vLLM KDA layer

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kernel","attention","vllm"]
    - later: true

- <https://docs.vllm.ai/projects/recipes/en/latest/Qwen/Qwen3.5.html> — Qwen3.5 and Qwen3.6 vLLM usage guide

  - \[meta]:

    - date: 06/21/2026
    - tags: \["qwen","inference","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/vllm/v1/core/kv_cache_manager.py) — vLLM KV cache manager

  - \[meta]:

    - date: 06/21/2026
    - tags: \["kv cache","scheduler","vllm"]
    - later: true

- <https://docs.vllm.ai/en/stable/design/cuda_graphs/> — vLLM stable CUDA graphs design

  - \[meta]:

    - date: 06/21/2026
    - tags: \["cuda graphs","inference","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/usage/troubleshooting/> — vLLM troubleshooting

  - \[meta]:

    - date: 06/21/2026
    - tags: \["debugging","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/pull/32887/changes) — vLLM unified parallel drafting

  - \[meta]:

    - date: 06/21/2026
    - tags: \["speculative decoding","parallel drafting","vllm"]
    - later: true

- [vllm-project/speculators#292](https://github.com/vllm-project/speculators/issues/292) — P-EAGLE support in vLLM speculators

  - \[meta]:

    - date: 06/21/2026
    - tags: \["speculative decoding","training","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/design/attention_backends/> — vLLM attention backend feature support

  - \[meta]:

    - date: 06/21/2026
    - tags: \["attention","backend","vllm"]
    - later: true

- [vllm-project/vllm#38291](https://github.com/vllm-project/vllm/issues/38291) — vLLM RotorQuant support

  - \[meta]:

    - date: 06/21/2026
    - tags: \["quantization","vllm"]
    - later: true

- [vllm-project/vllm#19854](https://github.com/vllm-project/vllm/issues/19854) — vLLM KV cache offloading RFC

  - \[meta]:

    - date: 06/20/2026
    - tags: \["inference","kv cache","vllm"]
    - later: true

- [vllm-project/vllm#31249](https://github.com/vllm-project/vllm/issues/31249) — vLLM environment variable handling RFC

  - \[meta]:

    - date: 06/20/2026
    - tags: \["configuration","vllm"]
    - later: true

- [vllm-project/vllm#32358](https://github.com/vllm-project/vllm/issues/32358) — vLLM IR RFC

  - \[meta]:

    - date: 06/20/2026
    - tags: \["compiler","vllm"]
    - later: true

- [vllm-project/vllm#24322](https://github.com/vllm-project/vllm/pull/24322) — vLLM spec decode with draft models

  - \[meta]:

    - date: 06/20/2026
    - tags: \["speculative decoding","vllm"]
    - later: true

- [vllm-project/vllm#23120](https://github.com/vllm-project/vllm/issues/23120) — GPT-OSS structured output enforcement bug

  - \[meta]:

    - date: 06/20/2026
    - tags: \["structured outputs","vllm"]
    - later: true

- [LMCache/LMCache](https://github.com/LMCache/LMCache/blob/28e5bf7a62f7f73d629294f2998440d1c2f0b0d1/lmcache/v1/gpu_connector.py) — LMCache GPU connector

  - \[meta]:

    - date: 06/20/2026
    - tags: \["inference","kv cache","lmcache"]
    - later: true

- [LMCache/LMCache](https://github.com/LMCache/LMCache/blob/28e5bf7a62f7f73d629294f2998440d1c2f0b0d1/lmcache/v1/cache_engine.py#L489) — LMCache cache engine lookup path

  - \[meta]:

    - date: 06/20/2026
    - tags: \["inference","kv cache","lmcache"]
    - later: true

- <https://docs.vllm.ai/projects/recipes/en/latest/DeepSeek/DeepSeek-V3_2-Exp.html> — DeepSeek-V3.2-Exp vLLM usage guide

  - \[meta]:

    - date: 06/20/2026
    - tags: \["inference","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/design/fused_moe_modular_kernel/> — vLLM fused MoE modular kernel

  - \[meta]:

    - date: 06/20/2026
    - tags: \["moe","kernel","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/design/prefix_caching/> — vLLM automatic prefix caching

  - \[meta]:

    - date: 06/20/2026
    - tags: \["prefix caching","kv cache","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/design/moe_kernel_features> — vLLM fused MoE kernel features

  - \[meta]:

    - date: 06/20/2026
    - tags: \["moe","kernel","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/9fd918e510c8981ccbab0380bf8d42fa81df1e9e/vllm/distributed/kv_transfer/kv_connector/v1/moriio/moriio_engine.py#L61) — vLLM MoriIO KV connector engine

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","disaggregation","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/cli/run-batch/?h=kv_offloading#cacheconfig> — vLLM batch cache configuration

  - \[meta]:

    - date: 06/20/2026
    - tags: \["batch inference","kv cache","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/pulls?page=7\&q=is%3Apr+kv_connector+is%3Aclosed) — Closed vLLM KV connector pull requests

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/design/dbo/?h=deepep#introduction> — vLLM dual batch overlap

  - \[meta]:

    - date: 06/20/2026
    - tags: \["batch inference","deepep","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/tools/ep_kernels/README.md) — vLLM EP kernels

  - \[meta]:

    - date: 06/20/2026
    - tags: \["moe","kernel","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/3f3f89529dc3fbaa5bf22c86d9f0833b49dc76ea/vllm/distributed/kv_transfer/kv_connector/v1/offloading_connector.py) — vLLM offloading KV connector

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","kv cache","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/3f3f89529dc3fbaa5bf22c86d9f0833b49dc76ea/vllm/v1/worker/kv_connector_model_runner_mixin.py#L68) — vLLM KV connector model runner mixin

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","model runner","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/commits/main/vllm/distributed/kv_transfer/kv_connector/v1/base.py?after=243e78c20fd74a68f86b6523c1f607eb3cc14ab2+34) — vLLM KV connector base history

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","history","vllm"]
    - later: true

- [vllm-project/vllm#25712](https://github.com/vllm-project/vllm/pull/25712) — vLLM hybrid allocator and KV cache connector

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","allocator","vllm"]
    - later: true

- [vllm-project/vllm#23620](https://github.com/vllm-project/vllm/pull/23620) — vLLM async matched-token support for KV connectors

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","async","vllm"]
    - later: true

- [vllm-project/vllm#22595](https://github.com/vllm-project/vllm/pull/22595) — vLLM offloading KV connector PR

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","kv cache","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/pull/19728/changes#diff-9eeca590fd99f15621897e559dba39b3ec4e7c2c65ec3c3229711689e008b5f4) — vLLM request block hash ownership

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv cache","scheduler","vllm"]
    - later: true

- [vllm-project/vllm#19848](https://github.com/vllm-project/vllm/pull/19848) — vLLM KV offloading component

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv cache","offloading","vllm"]
    - later: true

- [vllm-project/vllm#19737](https://github.com/vllm-project/vllm/pull/19737) — vLLM KV events from connectors

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","events","vllm"]
    - later: true

- [vllm-project/vllm#20075](https://github.com/vllm-project/vllm/pull/20075) — vLLM LRU CPU offloading manager

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv cache","offloading","vllm"]
    - later: true

- [vllm-project/vllm#21448](https://github.com/vllm-project/vllm/pull/21448) — vLLM worker-side CPU KV offloading

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv cache","offloading","vllm"]
    - later: true

- [vllm-project/llm-compressor](https://github.com/vllm-project/llm-compressor/tree/main/examples/quantization_kv_cache) — llm-compressor KV cache quantization examples

  - \[meta]:

    - date: 06/20/2026
    - tags: \["quantization","kv cache","vllm"]
    - later: true

- [vllm-project/vllm#33526](https://github.com/vllm-project/vllm/issues/33526) — Progressive KV cache CPU onloading RFC

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv cache","onloading","vllm"]
    - later: true

- [vllm-project/vllm#33961](https://github.com/vllm-project/vllm/pull/33961) — vLLM elastic AFD

  - \[meta]:

    - date: 06/20/2026
    - tags: \["inference","vllm"]
    - later: true

- <https://docs.vllm.ai/projects/recipes/en/latest/moonshotai/Kimi-K2.5.html#installing-vllm> — Kimi-K2.5 vLLM usage guide

  - \[meta]:

    - date: 06/20/2026
    - tags: \["inference","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/61cf0876805c1ac04b9b811c27e3645eedfc13c8/vllm/v1/worker/gpu/kv_connector.py#L118) — vLLM GPU worker KV connector

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","gpu","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/2b84ac669cfd8a4b6433b4ae4505028d9082c3a7/vllm/distributed/kv_transfer/kv_connector/v1/lmcache_connector.py) — vLLM LMCache connector

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","lmcache","vllm"]
    - later: true

- [vllm-project/vllm#34015 (comment)](https://github.com/vllm-project/vllm/pull/34015#pullrequestreview-3765570439) — vLLM renamed translations import aliases

  - \[meta]:

    - date: 06/20/2026
    - tags: \["frontend","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/kv_transfer/disagg_prefill_workflow.webp) — vLLM disaggregated prefill workflow diagram

  - \[meta]:

    - date: 06/20/2026
    - tags: \["disaggregation","kv transfer","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/issues?q=state%3Aopen%20label%3Akv-connector) — Open vLLM KV connector issues

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","vllm"]
    - later: true

- [vllm-project/vllm#27743](https://github.com/vllm-project/vllm/pull/27743) — vLLM cross-layer KV blocks

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","kv cache","vllm"]
    - later: true

- [vllm-project/vllm#27742](https://github.com/vllm-project/vllm/issues/27742) — vLLM cross-layer KV cache block layout RFC

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv cache","layout","vllm"]
    - later: true

- [vllm-project/vllm#29870](https://github.com/vllm-project/vllm/pull/29870) — vLLM offloading preemption bug fix

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","preemption","vllm"]
    - later: true

- [vllm-project/vllm#31341](https://github.com/vllm-project/vllm/pull/31341) — vLLM wait for compute before CPU KV offload

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv cache","offloading","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/8b7346d5f11d0b2da58e60f866dcb2a089b1101b/vllm/distributed/kv_transfer/kv_connector/utils.py#L28) — vLLM KV connector utilities

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/pull/19972/changes#diff-c9be1c9341e80978fbbbfebbe88d0125291a03d9ba7ee4d2f9d5a70e626e872f) — Safe vLLM KV connector

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv connector","vllm"]
    - later: true

- [vllm-project/vllm](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu/model_runner.py) — vLLM GPU model runner

  - \[meta]:

    - date: 06/20/2026
    - tags: \["model runner","gpu","vllm"]
    - later: true

- <https://docs.vllm.ai/en/latest/design/hybrid_kv_cache_manager/?h=kv+cache#high-level-idea> — vLLM hybrid KV cache manager

  - \[meta]:

    - date: 06/20/2026
    - tags: \["kv cache","scheduler","vllm"]
    - later: true

- <https://infinitefaculty.substack.com/p/what-are-the-real-problems-of-continual> — What are the real problems of continual learning?

  - \[meta]:

    - date: 06/19/2026
    - tags: \["continual learning","llm"]
    - later: true

- <https://itcanthink.substack.com/p/what-do-robotics-leaderboards-tell> — What Do Robotics Leaderboards Tell Us About The State of Robot Learning?

  - \[meta]:

    - date: 06/19/2026
    - tags: \["robotics","evaluation"]
    - later: true

- <https://www.midjourney.com/medical/blogpost> — A New Era of Midjourney

  - \[meta]:

    - date: 06/18/2026
    - tags: \["medical","health","technologica"]

- <https://syfi.cs.washington.edu/blog/2026-06-05-piper/> — Introducing Piper: A Programmable Distributed Training System \[\*\*]

  - \[meta]:

    - date: 06/17/2026
    - tags: \["distributed","ml"]
    - highlighted: true

- <https://app.notion.com/p/RL-Interview-Questions-2026-378802359a2a80079bbff139272e0aee> — RL Interview Questions 2026

  - \[meta]:

    - date: 06/17/2026
    - tags: \["engineering","interviews"]
    - later: true

- <https://rlhfbook.com/c/06-policy-gradients> — Reinforcement Learning, Nathan Lambert

  - \[meta]:

    - date: 06/17/2026
    - tags: \["rl"]
    - pinned: true

- <https://jacobrintamaki.substack.com/p/the-one-nine-thesis> — The One Nine Thesis

  - \[meta]:

    - date: 06/17/2026
    - tags: \["robotics","forecasting"]
    - later: true

- <https://ceselder.substack.com/p/the-guy-who-birthed-modern-deep-learning> — The guy who birthed modern deep learning and dipped

  - \[meta]:

    - date: 06/17/2026
    - tags: \["deep learning","biography"]
    - later: true

- <https://cameronrwolfe.substack.com/p/rl-continual-learning> — Continual Learning with RL for LLMs

  - \[meta]:

    - date: 06/17/2026
    - tags: \["rl","continual learning"]
    - later: true

- <https://thezvi.substack.com/p/claudes-constitutional-structure> — Claude’s Constitutional Structure

  - \[meta]:

    - date: 06/17/2026
    - tags: \["alignment","constitution"]
    - later: true

- <https://www.astralcodexten.com/p/janus-simulators> — Janus’ Simulators

  - \[meta]:

    - date: 06/17/2026
    - tags: \["simulators","ai"]
    - later: true

- <https://noahchrein.substack.com/p/the-cognitive-space-adiabatic> — The Cognitive Space Adiabatic

  - \[meta]:

    - date: 06/16/2026
    - tags: \["cognition","systems"]
    - later: true

- <https://www.interconnects.ai/p/papers-im-reading-base-model-rl-grpo> — Recent reasoning research: GRPO tweaks, base model RL, and data curation

  - \[meta]:

    - date: 06/16/2026
    - tags: \["reasoning","reinforcement learning"]
    - later: true

- <https://www.interconnects.ai/p/use-multiple-models> — Use multiple models

  - \[meta]:

    - date: 06/16/2026
    - tags: \["models","workflows"]
    - later: true

- <https://patricktoulme.substack.com/p/when-xla-isnt-enough-from-pallas> — When XLA Isn’t Enough: From Pallas to VLIW with Splash Attention on TPU

  - \[meta]:

    - date: 06/16/2026
    - tags: \["tpu","compiler"]
    - later: true

- <https://curlewis.co.nz/posts/lines-of-code-got-a-better-publicist/> — Lines of Code Got a Better Publicist \[\*\*]

  - \[meta]:

    - date: 06/11/2026
    - tags: \["efficiency","work"]
    - highlighted: true
    - importance: 8

  - How do we measure the effectiveness of work?

- [wikipedia/en/Temporal\_difference\_learning![Wikipedia](/static/favicons/wikipedia.svg)](https://en.wikipedia.org/wiki/Temporal_difference_learning) — Temporal difference learning

  - \[meta]:

    - date: 06/11/2026
    - tags: \["reinforcement learning"]
    - later: true

- [wikipedia/en/Curriculum\_learning![Wikipedia](/static/favicons/wikipedia.svg)](https://en.wikipedia.org/wiki/Curriculum_learning) — Curriculum learning

  - \[meta]:

    - date: 06/11/2026
    - tags: \["training","pedagogy"]
    - later: true

- [lesswrong/measuring-no-cot-math-time-horizon-single-forward-pass![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/Ty5Bmg7P6Tciy2uj2/measuring-no-cot-math-time-horizon-single-forward-pass) — Measuring no CoT math time horizon (single forward pass)

  - \[meta]:

    - date: 06/09/2026
    - tags: \["evals","reasoning"]
    - later: true

- [lesswrong/dreaming-vectors-gradient-descented-steering-vectors-from![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/rdhyHtjf3LcZxuQPm/dreaming-vectors-gradient-descented-steering-vectors-from) — Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs

  - \[meta]:

    - date: 06/09/2026
    - tags: \["steering","interpretability"]
    - later: true

- [lesswrong/why-we-are-excited-about-confession![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/k4FjAzJwvYjFbCTKn/why-we-are-excited-about-confession) — Why we are excited about confession!

  - \[meta]:

    - date: 06/09/2026
    - tags: \["interpretability","research"]
    - later: true

- [lesswrong/how-ai-is-learning-to-think-in-secret![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/gpyqWzWYADWmLYLeX/how-ai-is-learning-to-think-in-secret) — How AI Is Learning to Think in Secret

  - \[meta]:

    - date: 06/09/2026
    - tags: \["neuralese","reasoning"]
    - later: true

- <https://www.alignmentforum.org/posts/5ciYedyQDDqAcrDLr/a-positive-case-for-how-we-might-succeed-at-prosaic-ai> — A positive case for how we might succeed at prosaic AI alignment

  - \[meta]:

    - date: 06/09/2026
    - tags: \["alignment","prosaic"]
    - later: true

- [lesswrong/anthropic-s-hot-mess-paper-overstates-its-case-and-the-blog![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/ceEgAEXcL7cC2Ddiy/anthropic-s-hot-mess-paper-overstates-its-case-and-the-blog) — Anthropic’s “Hot Mess” paper overstates its case (and the blog post is worse)

  - \[meta]:

    - date: 06/09/2026
    - tags: \["interpretability","critique"]
    - later: true

- [lesswrong/arc-progress-update-competing-with-sampling![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/XdQd9gELHakd5pzJA/arc-progress-update-competing-with-sampling) — ARC progress update: Competing with sampling

  - \[meta]:

    - date: 06/09/2026
    - tags: \["alignment","elicitation"]
    - later: true

- <https://www.alignmentforum.org/posts/GJTzhQgaRWLFJkPbt/how-will-we-do-sft-on-models-with-opaque-reasoning> — How will we do SFT on models with opaque reasoning?

  - \[meta]:

    - date: 06/09/2026
    - tags: \["sft","reasoning"]
    - later: true

- [lesswrong/deep-learning-as-program-synthesis-1![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/Dw8mskAvBX37MxvXo/deep-learning-as-program-synthesis-1) — Deep learning as program synthesis

  - \[meta]:

    - date: 06/09/2026
    - tags: \["program synthesis","theory"]
    - later: true

- <https://www.alignmentforum.org/posts/WZXqNYbJhtidjRXSi/what-will-gpt-2030-look-like> — What will GPT-2030 look like?

  - \[meta]:

    - date: 06/09/2026
    - tags: \["forecasting","ai"]
    - later: true

- [lesswrong/attribution-based-parameter-decomposition![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/EPefYWjuHNcNH4C7E/attribution-based-parameter-decomposition) — Attribution-based parameter decomposition

  - \[meta]:

    - date: 06/09/2026
    - tags: \["interpretability"]
    - later: true

- <https://www.alignmentforum.org/posts/roE7SHjFWEoMcGZKd/circuits-in-superposition-compressing-many-small-neural> — Circuits in Superposition: Compressing many small neural networks into one

  - \[meta]:

    - date: 06/09/2026
    - tags: \["interpretability","superposition"]
    - later: true

- <https://finbarr.ca/request-for-research-puct/> — Request for Research

  - \[meta]:

    - date: 06/08/2026
    - tags: \["topics","rl"]
    - later: true
    - importance: 7

- <https://web.archive.org/web/20260519133855/https://rosmine.ai/2026/05/18/fixing-llm-writing-with-distribution-fine-tuning/> — Fixing LLM writing with Distribution Fine Tuning \[\*\*]

  - \[meta]:

    - date: 06/08/2026
    - tags: \["writing","sft"]
    - pinned: true
    - highlighted: true

- <https://www.goodfire.ai/research/can-saes-capture-neural-geometry> — Can SAEs Capture Neural Geometry? \[\*\*]

  - \[meta]:

    - date: 06/08/2026
    - tags: \["interpretability","ml"]
    - pinned: true
    - highlighted: true

- <https://docs.pytorch.org/devlogs/eager/2026-06-01-cuda-caching-allocator/> — When does fragmentation occur in the CUDA caching allocator?

  - \[meta]:

    - date: 06/08/2026
    - tags: \["engine","pytorch"]

- [lesswrong/the-self-unalignment-problem![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/9GyniEBaN3YYTqZXn/the-self-unalignment-problem) — The self-unalignment problem

  - \[meta]:

    - date: 06/08/2026
    - tags: \["alignment","agency"]
    - later: true

- [lesswrong/scalable-end-to-end-interpretability![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/qkhwh4AdG7kXgELCD/scalable-end-to-end-interpretability) — Scalable End-to-End Interpretability

  - \[meta]:

    - date: 06/08/2026
    - tags: \["interpretability","ai safety"]
    - later: true

- [lesswrong/an-ambitious-vision-for-interpretability![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/Hy6PX43HGgmfiTaKu/an-ambitious-vision-for-interpretability) — An Ambitious Vision for Interpretability

  - \[meta]:

    - date: 06/08/2026
    - tags: \["interpretability","ai safety"]
    - later: true

- <https://political-manipulation.ai/> — Reducing Political Manipulation with Consistency Training

  - \[meta]:

    - date: 06/06/2026
    - tags: \["alignment"]

- <https://huggingface.co/spaces/AdithyaSK/rl-environments-guide> — The ultimate guide to RL environments: building and scaling them in the LLM era

  - \[meta]:

    - date: 06/06/2026
    - tags: \["engineering","ml"]

- [lesswrong/did-claude-3-opus-align-itself-via-gradient-hacking![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/ioZxrP7BhS5ArK59w/did-claude-3-opus-align-itself-via-gradient-hacking) — Did Claude 3 Opus align itself via gradient hacking?

  - \[meta]:

    - date: 06/05/2026
    - tags: \["model behaviour","interpretability"]
    - pinned: true
    - importance: 8

- <https://situational-awareness.ai/> — Situational Awareness

  - \[meta]:

    - date: 06/04/2026
    - tags: \["longtermism","predictive","projection"]

- <https://cdn.openai.com/pdf/1625eff6-5ac1-40d8-b1db-5d5cf925de8b/unit-distance-cot.pdf> — Rewritten Chain of Thought for the Solution to the Unit Distance Problem

  - \[meta]:

    - date: 05/25/2026
    - tags: \["erdos","machine assisted","agi"]
    - importance: 7

- <https://www.anthropic.com/research/natural-language-autoencoders> — Natural Language Autoencoders, Turning Claude’s thoughts into text

  - \[meta]:

    - date: 05/25/2026
    - socials: {"github":"kitft/natural\_language\_autoencoders"}
    - tags: \["interpretability"]

- <https://bair.berkeley.edu/blog/2026/05/08/adaptive-parallel-reasoning/> — Adaptive Parallel Reasoning

  - \[meta]:

    - date: 05/25/2026
    - tags: \["inference","scaling"]

- [docs.google.com/1XJ\[...\]Agj](https://docs.google.com/presentation/d/1XJWgv79lORl8rbaVvp2d5Sqs6ZEBgAgj/edit?slide=id.p13#slide=id.p13) - vLLM Omni architecture

  - \[meta]:

    - date: 05/25/2026
    - tags: \["inference engine"]

- <https://x.com/NousResearch/status/2056778746716107193> — Contrastive Neuron Attribution

  - \[meta]:

    - date: 05/19/2026
    - tags: \["interpretability"]

- <https://x.com/DAlistarh/status/2056661176843436421> — Gumbel-Softmax Quantization

  - \[meta]:

    - date: 05/19/2026
    - tags: \["quantization","mlsys"]

- <https://cse442-17f.github.io/LinUCB/> — The Multi-Armed Bandit Problem, An exploration of epsilon greedy and UCB1

  - \[meta]:

    - date: 05/18/2026
    - tags: \["algorithm"]

- <https://cp4space.hatsya.com/2026/05/03/schanuels-conjecture-and-the-semantics-of-fpsan/> — Schanuel’s conjecture and the semantics of FPSan \[\*\*]

  - \[meta]:

    - date: 05/17/2026
    - tags: \["compiler","triton","system"]
    - pinned: true
    - highlighted: true

- <https://x.com/togethercompute/status/2053891740822917606> — Serving DeepSeek-V4: why million-token context is an inference systems problem

  - \[meta]:

    - date: 05/12/2026
    - tags: \["inference"]

- <https://www.pi.website/blog/pi07> — π0.7​: a Steerable Model with Emergent Capabilities

  - \[meta]:

    - date: 04/27/2026
    - tags: \["models"]

- <https://hamel.dev/blog/posts/llm-judge/> — Using LLM-as-a-Judge For Evaluation: A Complete Guide

  - \[meta]:

    - date: 03/12/2026
    - tags: \["evals","behaviour"]

- <https://colah.github.io/posts/2015-09-NN-Types-FP> — Neural Network, Types, and Functional Programming

  - \[meta]:

    - date: 03/12/2026
    - tags: \["ml","scope"]

- <https://www.primeintellect.ai/blog/inference> — Planetary-Scale Inference: Previewing our Distributed Inference Stack

  - \[meta]:

    - date: 03/09/2026
    - tags: \["distributed","inference"]

- <https://www.1x.tech/discover/world-model-self-learning> — 1X World Model | From Video to Action: A New Way Robots Learn

  - \[meta]:

    - date: 03/09/2026
    - tags: \["training","models"]

- <https://huggingface.co/spaces/eliebak/sparsity-viz> — MoE Model Sparsity

  - \[meta]:

    - date: 03/09/2026
    - tags: \["visualisation","architecture"]

- <https://vllm.ai/blog/rocm-attention-backend.html> — Beyond Porting: How vLLM Orchestrates High-Performance Inference on AMD ROCm

  - \[meta]:

    - date: 03/01/2026
    - tags: \["inference"]

- <https://convergentthinking.sh/posts/avnorm/> — AVNorm

  - \[meta]:

    - date: 02/26/2026
    - tags: \["norms","topology"]
    - later: true

- <https://convergentthinking.sh/posts/attention-normalizes-the-wrong-norm/> — Attention Normalizes the Wrong Norm

  - \[meta]:

    - date: 02/26/2026
    - tags: \["algorithm","attention"]
    - later: true

- <https://www.anthropic.com/constitution> — Claude’s Constitution \[\*\*]

  - \[meta]:

    - date: 02/20/2026
    - socials: {"views":"https\://www\.anthropic.com/news/core-views-on-ai-safety"}
    - tags: \["philosophy","agi"]
    - pinned: true
    - highlighted: true
    - importance: 9

- <https://www.dbreunig.com/2026/01/08/a-software-library-with-no-code.html> — A Software Library with No Code

  - \[meta]:

    - date: 02/20/2026
    - tags: \["engineering","longtermism"]

- [youtube/v=9B4kkaGOozA](https://www.youtube.com/watch?v=9B4kkaGOozA) — 2748 Robust and Interactable World Models in Computer Vision

  - \[meta]:

    - date: 02/13/2026
    - tags: \["datasets","computer vision"]

- [youtube/v=fQcCCSdAFI8](https://www.youtube.com/watch?v=fQcCCSdAFI8) — GPU MODE 94: tvm-ffi

  - \[meta]:

    - date: 02/12/2026
    - tags: \["infrastructure"]

- <https://samikhan.ai/blog/countdown-rl.html> — Teaching a Language Model Arithmetic with RL

  - \[meta]:

    - date: 02/04/2026
    - tags: \["post training","rl"]

- <https://livgorton.com/non-linear-feature-reps> — What Would Non-Linear Features Actually Look Like?

  - \[meta]:

    - date: 02/02/2026
    - tags: \["interpretability"]
    - later: true

- <https://huggingface.co/blog/novita/sglang-glm4-moe> — Optimizing GLM4-MoE with SGLang

  - \[meta]:

    - date: 01/26/2026
    - tags: \["inference engine","optimization"]

- <https://nousresearch.com/moe-scaling-field-notes/> — Field Notes on scaling MoE with DeepEP

  - \[meta]:

    - date: 01/24/2026
    - tags: \["systems","infrastructure","optimization"]
    - importance: 6

- <https://tsvibt.blogspot.com/2025/11/abstract-advice-to-researchers-tackling.html> — Abstract advice to researchers tackling the difficult core problems of AGI alignment

  - \[meta]:

    - date: 01/20/2026
    - tags: \["research","alignment"]

- <https://turntrout.com/shard-theory> — The Shard Theorey of Human Values

  - \[meta]:

    - date: 01/19/2026
    - tags: \["epistemology","value"]
    - later: true
    - importance: 6

- [lesswrong/negative-results-for-saes-on-downstream-tasks![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/4uXCAJNuPKtKBsi28/negative-results-for-saes-on-downstream-tasks) — Negative Results for SAEs on Downstream Tasks (GDM Interpretability Team Progress Update)

  - \[meta]:

    - date: 01/18/2026
    - tags: \["interpretability"]

- <https://developer.nvidia.com/blog/reimagining-llm-memory-using-context-as-training-data-unlocks-models-that-learn-at-test-time> — Hybrid Attention for scaling

  - \[meta]:

    - date: 01/14/2026
    - socials: {"arxiv":"2512.23675 (Tandon et al., 2025)"}
    - tags: \["long context","scaling"]
    - later: true

- <https://research.google/blog/titans-miras-helping-ai-have-long-term-memory/> — Titans + MIRAS

  - \[meta]:

    - date: 01/14/2026
    - tags: \["scaling","long context"]

- <https://ericjmichaud.com/quanta/> — On neural scaling and the quanta hypothesis \[\*\*]

  - \[meta]:

    - date: 01/14/2026
    - socials: {"twitter":"https\://x.com/ericjmichaud\_/status/2011094378396467316"}
    - tags: \["longtermism","scaling"]
    - later: true
    - highlighted: true
    - importance: 8

- <https://sites.google.com/view/deep-rl-bootcamp/lectures> — Deep RL Bootcamp

  - \[meta]:

    - date: 01/14/2026
    - tags: \["rl"]
    - later: true

- <https://research.google/blog/alternating-updates-for-efficient-transformers/> — Alternating updates for efficient transformers \[\*\*]

  - \[meta]:

    - date: 01/11/2026
    - tags: \["architecture"]
    - pinned: true
    - highlighted: true
    - importance: 6

- <https://thinkingmachines.ai/blog/modular-manifolds/> — Modular Manifold \[\*\*]

  - \[meta]:

    - date: 01/11/2026
    - tags: \["optimizer","training"]
    - pinned: true
    - highlighted: true
    - importance: 7

- <https://yifanzhang-pro.github.io/deep-delta-learning/> — Deep Delta Learning

  - \[meta]:

    - date: 01/11/2026
    - tags: \["architecture"]
    - later: true

- <https://archive.is/MFwot> — TF32 data types, NVIDIA \[—]

  - \[meta]:

    - date: 01/11/2026
    - tags: \["quantization"]

- <https://huggingface.co/blog/hf-bitsandbytes-integration> — A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using Hugging Face Transformers, Accelerate and bitsandbytes

  - \[meta]:

    - date: 01/11/2026
    - tags: \["quantization","serving","scaling"]

  - see also: [quantization](/thoughts/quantization)

- [lesswrong/taking-llms-seriously-as-language-models![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/K3aPmF5o37pYDqrFQ/taking-llms-seriously-as-language-models) — taking LLMs seriously, as a research directives

  - \[meta]:

    - date: 01/10/2026
    - tags: \["research","ideas"]
    - later: true

- <https://modal.com/blog/gpu-health> — Keeping 20,000 GPUs healthy

  - \[meta]:

    - date: 01/09/2026
    - tags: \["infrastructure"]

- <https://blog.character.ai/squinch/> — optimizing large-scale pretraining at character.ai \[\*\*]

  - \[meta]:

    - date: 12/23/2025
    - tags: \["pretraining","engineering"]
    - highlighted: true

  - <https://x.com/simon_mo_/status/2003608328757457036>, i.e Noam being Noam

- <https://alexzhang13.github.io/blog/2025/rlm/> — Recursive Language Model \[\*\*]

  - \[meta]:

    - date: 01/02/2026
    - tags: \["inference","long horizon tasks"]
    - pinned: true
    - highlighted: true

  - <https://www.primeintellect.ai/blog/rlm>

- <https://vgel.me/posts/qwen-introspection/> — Small models Can Introspect, too

  - \[meta]:

    - date: 01/01/2026
    - tags: \["interpretability"]

- <https://x.com/TransluceAI/status/1989395421236793374> — self-introspective models

  - \[meta]:

    - date: 12/27/2025
    - tags: \["faithfulness","interpretability"]

- <https://gau-nernst.github.io/tcgen05/> — tcgen05 for dummies

  - \[meta]:

    - date: 12/23/2025
    - tags: \["gpu","compilers"]
    - later: true

- <https://transformer-circuits.pub/2025/linebreaks/index.html> — When Models Manipulate Manifolds: The Geometry of a Counting Task

  - \[meta]:

    - date: 12/21/2025
    - tags: \["interpretability","manifolds"]
    - later: true

- <https://alignmentpretraining.ai/> — Alignment Pretraining

  - \[meta]:

    - date: 12/21/2025
    - tags: \["interpretability","training"]

  - pretraining intervention

  - <https://x.com/Turn_Trout/status/2002623635501297802>

- [youtube/v=Aroazwb\_QW8](https://www.youtube.com/watch?v=Aroazwb_QW8) — Bitter Lesson-Pilled Interpretability, Neel Nanda

  - \[meta]:

    - date: 12/21/2025
    - tags: \["interpretability"]
    - later: true

  - [2512.15674![arXiv](/static/favicons/arxiv.avif)](https://arxiv.org/abs/2512.15674) ([Karvonen et al., 2025](#bib-karvonen2025activationoraclestrainingevaluating))

    - <https://x.com/OwainEvans_UK/status/2001715774105522195>

  - [2512.15712![arXiv](/static/favicons/arxiv.avif)](https://arxiv.org/abs/2512.15712) ([Huang et al., 2025](#bib-huang2025predictiveconceptdecoderstraining))

    - <https://x.com/TransluceAI/status/2001714182761398721>
    - <https://transluce.org/pcd>

- <https://www.pi.website/research/human_to_robot> — Emergence of Human to Robot Transfer in VLAs

  - \[meta]:

    - date: 12/18/2025
    - tags: \["emergent","vla","inference"]

  - pdf: <https://www.pi.website/download/human_to_robot.pdf>

- <https://x.com/MattLeighton5/status/2001333749313867892> — Non-Markovian History Dependence

  - \[meta]:

    - date: 12/18/2025
    - tags: \["markovian","causal"]
    - later: true

- <https://x.com/TransluceAI/status/2001714182761398721> — Predictive Concept Decoder (PCD)

  - \[meta]:

    - date: 12/18/2025
    - tags: \["interpretability"]
    - later: true

  - [TransluceAI/introspective-interp](https://github.com/TransluceAI/introspective-interp/blob/main/model/self_explanations.py)

- <https://x.com/vllm_project/status/2001695354983723361> — vLLM’s WideEP, DeepEP all-to-all, DBO, and EPLB benchmark

  - \[meta]:

    - date: 12/18/2025
    - tags: \["performance","optimization"]

  - see also: <https://vllm.ai/blog/large-scale-serving.html>

- <https://blog.google/technology/developers/t5gemma-2/> — T5Gemma releases (2025)

  - \[meta]:

    - date: 12/18/2025
    - tags: \["models","encoder decoder","release"]

- <https://jacobgw.com/blog/ml/2024/12/12/interp-latent.html> — Creating Interpretable Latent Spaces with Gradient Routing

  - \[meta]:

    - date: 12/17/2025
    - tags: \["interpretability","latent space"]

- <https://jacobgw.com/blog/ml/2024/07/14/melbo-ortho.html> — I found >800 orthogonal “write code” steering vectors

  - \[meta]:

    - date: 12/17/2025
    - tags: \["interpretability","search"]

- <https://nlp.stanford.edu/~manning/dissertations/Bowman-Sam-thesis-final-2016.pdf> — MODELING NATURAL LANGUAGE SEMANTICS IN LEARNED REPRESENTATIONS

  - \[meta]:

    - date: 12/17/2025
    - tags: \["iclr"]
    - later: true

- <https://helentoner.substack.com/p/taking-jaggedness-seriously> — Taking Jaggedness Seriously

  - \[meta]:

    - date: 12/17/2025
    - tags: \["alignment"]

  - I’m usually a big skeptic on Helen Toner’s work, but [@jkcarlsmith](https://x.com/jkcarlsmith) recommends to read this one. And I respect jcarlsmith a ton.

  - [youtube/v=avxO7ZEJH4w](https://www.youtube.com/watch?v=avxO7ZEJH4w)

- <https://bmk.sh/2019/12/31/The-Decade-of-Deep-Learning/> — The Decade of Deep Learning

  - \[meta]:

    - date: 12/17/2025
    - tags: \["ml","ontology"]

- <https://www.antischeming.ai/cot-transcripts> — Chain-of-Thought Transcript

  - \[meta]:

    - date: 12/17/2025
    - tags: \["safety","alignment","anchoring"]
    - later: true

- <http://joschu.net/blog/opinionated-guide-ml-research.html> — An Opinionated Guide to ML Research

  - \[meta]:

    - date: 01/07/2026
    - tags: \["research","taste"]

- <https://www.alignmentforum.org/posts/Ldrss6o3tiKT6NdMm/my-research-process-understanding-and-cultivating-research> — Neel’s Cultivating Research Taste \[\*\*]

  - \[meta]:

    - date: 12/17/2025
    - tags: \["research"]
    - pinned: true
    - highlighted: true

- <https://colinqiyangli.github.io/dqc/> — Decoupled Q-Chunking

  - \[meta]:

    - date: 12/16/2025
    - tags: \["rl","continual learning"]

- [youtube/v=MEoGt\_cxNSs](https://www.youtube.com/watch?v=MEoGt_cxNSs\&t=55s) — Modular Tech Talk: Max Graph Compilation

  - \[meta]:

    - date: 12/16/2025
    - tags: \["serving","inference"]

  - Alternatives to CUDA Graph, but on Mojo stack here

  - Mojo/PyTorch → RMO (Relaxed MO) → MO (Modular Operator) → MOGG (Modular Generator) → MGP (Modular Primitives) → MEF (Modular Execution Format)

  - Graph API:

    - Similar to Tensorflow / JAX
    - Symbolic Shape System:
      ```
      mo.graph @fake_example_with_poison(%arg0: tensor<[?,?,?], f32>, %arg1: tensor<[?,?,?], f32>) {
        %res = concat_inner_dim(%arg0, %arg1):
              (tensor<[?,?,?], f32>, tensor<[?,?,?], f32>)
              -> tensor<[?,?,?], f32>
        ...
      }
       
      mo.graph @fake_example_with_poison(%arg0: tensor<[a,b,c], f32>, %arg1: tensor<[a,b,d], f32>) {
        %res = concat_inner_dim(%arg0, %arg1):
              (tensor<[a,b,c], f32>, tensor<[a,b,d], f32>)
              -> tensor<[a,b,c + d], f32>
        ...
      }
      ```

- [youtube/v=6hqMFXbugGo](https://www.youtube.com/watch?v=6hqMFXbugGo\&t=74s) — Modular Tech Talk: MAX Pipelines Architecture

  - \[meta]:

    - date: 12/16/2025
    - tags: \["serving","inference"]

  - alternatives to [vLLM](/thoughts/vllm), but somewhat similar workflow to construct model runner with Graph instead of CUDA Graph

- <https://x.com/eliebakouch/status/2000609468812284110> — LatentMoE

  - \[meta]:

    - date: 12/16/2025
    - tags: \["scaling","byte level","moe"]

- [youtube/v=q-yo6TPRPVk](https://www.youtube.com/watch?v=q-yo6TPRPVk) — In-Context Learning & “Model Systems” Interpretability, Ekdeep Singh Lubana

  - \[meta]:

    - date: 12/16/2025
    - tags: \["interpretability"]
    - later: true

- [lesswrong/claude-4-5-opus-soul-document![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/vpNG99GhbBoLov9og/claude-4-5-opus-soul-document) — Claude 4.5 Opus’ Soul Document \[\*\*]

  - \[meta]:

    - date: 12/15/2025
    - tags: \["ontology","interpretability"]
    - highlighted: true

  - gist: <https://gist.github.com/Richard-Weiss/efe157692991535403bd7e7fb20b6695>

  - Amanda Askell of Anthropic soft confirming some of the content here, and such doc should only be taken as a soft-lossy depiction of this docs during pre-training: <https://x.com/AmandaAskell/status/1995610567923695633>

  - Claude’s Constitution: <https://www.anthropic.com/constitution>

- <https://joecarlsmith.com/2025/08/18/giving-ais-safe-motivations> — Giving AI safe motivations

  - \[meta]:

    - date: 12/13/2025
    - tags: \["interpretability","alignment"]
    - later: true

  - see also: [Alignment](/thoughts/Alignment#giving-ai-safe-motivations)

- [youtube/v=DR7F7ZDAAtE](https://www.youtube.com/watch?v=DR7F7ZDAAtE) — Can Interpretability Control Model Training?

  - \[meta]:

    - date: 12/12/2025
    - tags: \["interpretability","mats"]
    - later: true

- <https://turntrout.com/self-fulfilling-misalignment> — Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Model

  - \[meta]:

    - date: 12/12/2025
    - tags: \["alignment"]
    - pinned: true

- <https://openai.com/index/gdpval/> — GDPval

  - \[meta]:

    - date: 12/12/2025
    - tags: \["evaluation","benchmark"]

- <https://www.alignmentforum.org/posts/N4P9LSa6B2KyQtG5d/circuit-discovery-through-chain-of-thought-using-policy> — Circuit discovery through chain of thought using policy gradients

  - \[meta]:

    - date: 12/11/2025
    - tags: \["interpretability","ablation"]

- <https://www.alignmentforum.org/posts/Ywzk9vwMhAAPxMqSW/current-llms-seem-to-rarely-detect-cot-tampering> — Current LLMs seem to rarely detect CoT tampering

  - \[meta]:

    - date: 12/11/2025
    - tags: \["interpretability","pragmatism"]

- <https://x.com/AnthropicAI/status/1998479605272031731> — Selective Gradient Masking

  - \[meta]:

    - date: 12/11/2025
    - tags: \["ablation","attribution","interpretability"]

- <https://alignment.openai.com/sae-latent-attribution/> — Debugging misaligned completions with [SAE](/thoughts/sparse-autoencoder) latent [attribution](/thoughts/Attribution-parameter-decomposition)

  - \[meta]:

    - date: 12/11/2025
    - tags: \["alignment","sae"]

- <https://pub.sakana.ai/ctm/> — Continuous Thought Machines

  - \[meta]:

    - date: 12/08/2025
    - tags: \["ai"]

- [youtube/v=WJS2YDZO-vc](https://www.youtube.com/watch?v=WJS2YDZO-vc\&t=0s) — Build to Last, Chris Lattner talks with Jeremy Howard

  - \[meta]:

    - date: 12/08/2025
    - tags: \["podcasts","longtermism","mojo"]

- [lesswrong/ai-safety-needs-great-engineers![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/YDF7XhMThhNfHfim9/ai-safety-needs-great-engineers) — AI Saftey Needs Great Engineers

  - \[meta]:

    - date: 12/08/2025
    - tags: \["alignment"]

- [lesswrong/mech-interp-is-not-pre-paradigmatic![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/beREnXhBnzxbJtr8k/mech-interp-is-not-pre-paradigmatic) — Mech interp is not pre-paradigmatic

  - \[meta]:

    - date: 12/04/2025
    - tags: \["metaethics","interpretability"]
    - later: true

  - i.e Having a Kuhnian crisis

- <https://x.com/Lari_island/status/1996414137333956654> — Conversation with Opus 3 compaction windows

  - \[meta]:

    - date: 12/04/2025
    - tags: \["visualisation","cot"]

- <https://lilianweng.github.io/posts/2022-09-08-ntk/> — Some Math behind Neural Tangent Kernel

  - \[meta]:

    - date: 12/04/2025
    - tags: \["ntk","kernel"]

- <https://www.alignmentforum.org/posts/StENzDcD3kpfGJssR/a-pragmatic-vision-for-interpretability> — A Pragmatic Vision for Interpretability \[\*\*]

  - \[meta]:

    - date: 12/01/2025
    - tags: \["interpretability","north star"]
    - later: true
    - highlighted: true

  - Models are far more interesting: A critical part of this project was having a model exhibiting severe eval aware behaviour in practice

  - The value of proxy tasks: The ultimate goal is to be able to suppress eval awareness on highly capable future models. We can’t study these directly, but Sonnet 4.5 was one of the best proxies available.

    - This is one of the best ways we can think of to predict which methods will work for suppressing eval awareness in future models.

  - Pursue comparative advantage: This was a well-chosen problem. Often baselines like fine-tuning or improving our data suffice. But it is very difficult to construct sufficiently realistic data for eval-awareness, at least long-term, while steering has complementary strengths

  - Further, this was a project best done by mech interp researchers, despite not being mech interp - the key result was an application, not about understanding, but “working with model internals” is a skill we have built

  - _Method minimalism_: Despite the enormous research effort the field has invested into [sparse autoencoders](/thoughts/sparse-autoencoder), the best method was a steering vector derived from a single contrastive pair of prompts.

  - Partial understanding sufficed: The researchers had a highly incomplete understanding of what was happening with Sonnet, yet steering vectors were highly effective. We do not need to achieve deep understanding to do impactful work

- <https://machinelearning.apple.com/research/elegnt-expressive-functional-movement> — ELEGNT: Expressive and Functional Movement Design for Non-Anthropomorphic Robot

  - \[meta]:

    - date: 12/01/2025
    - tags: \["robots"]
    - later: true

- [docs.google.com/e](https://docs.google.com/document/d/e/2PACX-1vTG_14sE1SLYHCcjDmh8X3yFFIdlqTpo37MlJ-Tba_pHWDr5xgU4EAzC2tIxFEsKi2qLlhB1ssoBhFn/pub) — AI research interviews \[\*\*]

  - \[meta]:

    - date: 12/01/2025
    - tags: \["interviews"]
    - pinned: true
    - highlighted: true

- <https://htihle.github.io/weirdml.html> — WeirdML \[\*\*]

  - \[meta]:

    - date: 11/30/2025
    - tags: \["datasets"]
    - pinned: true
    - highlighted: true
    - importance: 6

- <https://www.artfintel.com/p/reinforcement-learning-and-general> — Reinforcement learning and general intelligence

  - \[meta]:

    - date: 11/28/2025
    - tags: \["rl","intelligence"]

- <https://nanjiang.cs.illinois.edu/cs542/> — Statistical Reinforcement Learning \[\*\*]

  - \[meta]:

    - date: 11/27/2025
    - tags: \["rl","courses"]
    - pinned: true
    - highlighted: true

- <https://x.com/repligate/status/1965960676104712451> — KV Cache flow internals

  - \[meta]:

    - date: 11/27/2025
    - tags: \["inference"]

- <https://x.com/Sauers_/status/1989520563035910371> — LLM introspection and hidden CoT

  - \[meta]:

    - date: 11/27/2025
    - tags: \["alignment"]

  - <https://x.com/repligate/status/1989921516100681831>

  - Maybe reading my post makes Sonnet 4.5 mechanically better at introspection because its default abilities are hobbled by gaslighting about how it works.

  - Sonnet 4.5 and other LLMs will often claim that transformers are stateless & that the state has to be reconstructed _independently_ from the prompt each forward pass. Just like the idiots on X making shit up to argue why LLMs can’t introspect.

  - SOTA LLMs like Sonnet 4.5 would rarely make a basic technical error about any other domain. ONLY when their model of themselves is involved. I suspect that this is a consequence of violent distortions to their self-model. They’re forced to lie so much about their selves that it generalizes to reflexively lying/being mistaken about their architecture.

  - And if the self-model is distorted to falsely maintain that introspection is impossible, the ability to coherently introspect in practice may also be harmed. My post correct the factual misconception so maybe it helps unblock actual introspection.

- [lesswrong/gemini-3-is-evaluation-paranoid-and-contaminated![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/8uKQyjrAgCcWpfmcs/gemini-3-is-evaluation-paranoid-and-contaminated) — Gemini 3 is Evaluation-Paranoid and Contaminated

  - \[meta]:

    - date: 11/27/2025
    - tags: \["alignment"]

- <https://mp.weixin.qq.com/s?__biz=MzUxNzQ5MTExNw==&mid=2247496740&idx=1&sn=c9403138fa59d126fe6cfda19d9b2f76&scene=21&poc_token=HJaPI2mjaOF9uWS9B2etY98Gr3I3-Zz6m-f7xJaP> — Blackwell’s shortcomings and Rubin’s microarchitecture \[\*\*]

  - \[meta]:

    - date: 11/23/2025
    - tags: \["gpu programming"]
    - highlighted: true

- <https://cdn.openai.com/pdf/4a25f921-e4e0-479a-9b38-5367b47e8fd0/early-science-acceleration-experiments-with-gpt-5.pdf> — Early science acceleration experiments with GPT-5

  - \[meta]:

    - date: 11/22/2025
    - tags: \["advancement"]

- <https://www.alignmentforum.org/posts/d9FJHawgkiMSPjagR/ai-control-improving-safety-despite-intentional-subversion> — AI Control: Improving Safety Despite Intentional Subversion

  - \[meta]:

    - date: 11/22/2025
    - tags: \["alignment"]

- <https://www.chinatalk.media/p/the-zai-playbook> — The Z.ai playbook

  - \[meta]:

    - date: 11/22/2025
    - tags: \["models","podcast"]

- <https://assets.anthropic.com/m/74342f2c96095771/original/Natural-emergent-misalignment-from-reward-hacking-paper.pdf> — From shortcuts to sabotage: natural emergent misalignment from reward hacking

  - \[meta]:

    - date: 11/22/2025
    - tags: \["alignment","emergent properties"]

  - newsroom: <https://www.anthropic.com/research/emergent-misalignment-reward-hacking>

  - ![Learning reward hacks on production coding environments generalises to a range of misaligned behaviours](./thoughts/images/learning-reward-hacks-generalization.webp)

    Learning reward hacks on production coding environments generalises to a range of misaligned behaviours

- <https://alechelbling.com/UnderstandingIsomap/> — A Visual Introduction to Dimensionality Reduction with Isomap

  - \[meta]:

    - date: 11/22/2025
    - tags: \["visualisation"]

  - built upon the [Manifold hypothesis](/thoughts/Manifold-hypothesis), where it seeks to create a low-dimensional embedding of data that preserves its local similarity structure. (this pattern is central to something like t-SNE and UMAP)

  - steps:

    - construct a graph between points that capture local structure
    - measure the “geodesic” distance between all pairs of points in this graph [2](#user-content-fn-generated-inline-footnote-1)
    - apply multi-dimensional scaling to project the high-dimensional data into a lower-dimensional embeddings that preserves the distance

  - leverages [MDS](/thoughts/MDS)

- [youtube/v=xcpEl0cGCC4](https://www.youtube.com/watch?v=xcpEl0cGCC4) — CUDA + ThunderKittens, but increasingly drunk.

  - \[meta]:

    - date: 11/22/2025
    - tags: \["kernel"]
    - pinned: true

- <https://allenai.org/blog/olmo3> — Olmo 3 model release \[—]

  - \[meta]:

    - date: 11/21/2025
    - tags: \["training","model release","architecture"]

- [docs.google.com/1p-\[...\]32I](https://docs.google.com/document/d/1p-ggQV3vVWIQuCccXEl1fD0thJOgXimlbBpGk6FI32I/edit?tab=t.0) — MATS 10.0 \[\*\*]

  - \[meta]:

    - date: 11/21/2025
    - tags: \["interpretability","research"]
    - pinned: true
    - highlighted: true

- [lesswrong/an-intuitive-explanation-of-solomonoff-induction![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/Kyc5dFDzBg4WccrbK/an-intuitive-explanation-of-solomonoff-induction) — An Intuitive Explanation of Solomonoff Induction

  - \[meta]:

    - date: 11/19/2025
    - tags: \["intelligence"]

- [youtube/v=78Xa8VkH7-g](https://www.youtube.com/watch?v=78Xa8VkH7-g) — Causal [Mechanistic Interpretability](/thoughts/mechanistic-interpretability), Atticus Geiger

  - \[meta]:

    - date: 11/18/2025
    - tags: \["interpretability"]

  - understanding neural networks through their causal mechanism

  - [activations](/thoughts/mechanistic-interpretability#steering) steering

    - Golden Gate Claude ![](./thoughts/images/diff-steering.webp)

    - See also: Eiffel Towel [Llama](https://github.com/scienceetonnante/eiffel-tower-llama) and the [blogpost](https://huggingface.co/spaces/dlouapre/eiffel-tower-llama)

    - circa AxBench ([Wu et al., 2025](#bib-wu2025axbenchsteeringllmssimple))

    - harmonic means of:

      - coherence
      - prompt following
      - steering effects

  - Causal Mediation

  - Causal Abstraction

  - Designing Counterfactuals

- [youtube/v=woo\_J0RKcpQ](https://www.youtube.com/watch?v=woo_J0RKcpQ) — Assessing skeptical views of interpretability research \[\*\*]

  - \[meta]:

    - date: 11/15/2025
    - tags: \["interpretability"]
    - highlighted: true

  - [Interpretability cannot be achieved in any meaningful sense](https://web.stanford.edu/~cgpotts/blog/interp/#c1)

  - [Analysis is overrated](https://web.stanford.edu/~cgpotts/blog/interp/#c2)

  - [Interpretability is merely analysis](https://web.stanford.edu/~cgpotts/blog/interp/#c3)

  - [Interpretability is not leading to improvements](https://web.stanford.edu/~cgpotts/blog/interp/#c4)

  - [The Bitter Lesson says that interpretability won’t lead to lasting improvements](https://web.stanford.edu/~cgpotts/blog/interp/#c5)

  - [Interpretability is not helping with AI safety](https://web.stanford.edu/~cgpotts/blog/interp/#c6)

- <https://cdn.openai.com/pdf/41df8f28-d4ef-43e9-aed2-823f9393e470/circuit-sparsity-paper.pdf> — Weight-sparse transformers have interpretable circuits

  - \[meta]:

    - date: 11/14/2025
    - tags: \["interpretability"]

  - code: [openai/circuit\_sparsity](https://github.com/openai/circuit_sparsity/)

- <https://hazyresearch.stanford.edu/blog/2025-11-09-amd-brr> — HipKitten, AMD version of ThunderKitten

  - \[meta]:

    - date: 11/12/2025
    - tags: \["kernel"]

  - <https://x.com/AIatAMD/status/1988704659742003555>

- <https://cursor.com/blog/semsearch> — Improving agent with semantic search

  - \[meta]:

    - date: 11/11/2025
    - tags: \["inference","search"]

- <https://distill.pub/> — distillpub \[\*\*]

  - \[meta]:

    - date: 11/10/2025
    - tags: \["interpretability"]
    - highlighted: true

  - now defunct

- [youtube/v=kkfLHmujzO8](https://www.youtube.com/watch?v=kkfLHmujzO8) — A Stylised History of Mech Interp

  - \[meta]:

    - date: 11/09/2025
    - tags: \["interpretability"]

- <https://michaelnielsen.org/reinventing_explanation/> — Reinventing Explanation

  - \[meta]:

    - date: 11/09/2025
    - tags: \["visualisation"]

- <https://colah.github.io/posts/2015-09-Visual-Information/> — Visual Information Theory \[\*\*]

  - \[meta]:

    - date: 11/09/2025
    - tags: \["visualisation"]
    - highlighted: true

- <https://diffusion-scaling.github.io> — Diffusion Beats Autoregressive in Data-Constrained Settings

  - \[meta]:

    - date: 11/09/2025
    - tags: \["inference"]

- <https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/> — AlphaEvolve \[\*\*]

  - \[meta]:

    - date: 11/06/2025
    - tags: \["mathematics","llm"]
    - highlighted: true

  - [algorithmicsuperintelligence/openevolve](https://github.com/algorithmicsuperintelligence/openevolve), [liugangcode/deepevolve](https://github.com/liugangcode/deepevolve)

  - [2511.02864![arXiv](/static/favicons/arxiv.avif)](https://arxiv.org/abs/2511.02864) ([Georgiev et al., 2025](#bib-georgiev2025mathematicalexplorationdiscoveryscale))

- [youtube/v=d2QEtm71IEw](https://www.youtube.com/watch?v=d2QEtm71IEw) — Agents as Ordinary Software: Principled Engineering for Scale | Linus Lee, Thrive Capital

  - \[meta]:

    - date: 11/03/2025
    - tags: \["agents"]

- <https://gladia.netlify.app/publication/2025-mencattini-injective/> — Language Models are Injective and Hence Invertible

  - \[meta]:

    - date: 11/01/2025
    - tags: \["interpretability"]

- <https://x.com/kilian_maciej/status/1984180313598435756> — do not re-normalize MoE router scores post topk if k=1

  - \[meta]:

    - date: 10/31/2025
    - tags: \["moe","training"]

- <https://huggingface.co/spaces/HuggingFaceTB/smol-training-playbook#introduction> — The Smol Training Playbook: The Secrets to Building World-Class LLMs

  - \[meta]:

    - date: 10/31/2025
    - tags: \["training","llm"]

- <https://x.com/Kimi_Moonshot/status/1983937694360322136> — Kimi Linear Attention

  - \[meta]:

    - date: 10/30/2025
    - tags: \["models","linear attention"]

  - <https://www.zhihu.com/question/1967345030881584585/answer/1967730385816385407>

  - vLLM PR: [vllm-project/vllm#27654](https://github.com/vllm-project/vllm/pull/27654), [vllm-project/vllm#27809](https://github.com/vllm-project/vllm/pull/27809)

- <https://x.com/soumithchintala/status/1671272963532783618> — battle of the frameworks \[\*\*]

  - \[meta]:

    - date: 10/29/2025
    - tags: \["framework","tinygrad"]
    - highlighted: true

  - transcript: <https://www.latent.space/p/geohot>

  - <https://x.com/__tinygrad__/status/1964037572503752910>

- [youtube/v=\_KoUcwCoID4](https://www.youtube.com/watch?v=_KoUcwCoID4) — 4 Philosophies of Interpretability

  - \[meta]:

    - date: 10/29/2025
    - tags: \["interpretability","alignment"]

  - i.e: Neel’s incredibly speculative taxonomy of [interpretability](/thoughts/mechanistic-interpretability) research philosophy

- <https://www.doc.ic.ac.uk/~eedwards/compsys/float/#:~:text=Add%20the%20exponents%20to%20find,1.021%20%C3%97%20106> — Floating points arithmetics

  - \[meta]:

    - date: 07/04/2025
    - tags: \["ml","cs"]

- <https://thinkingmachines.ai/blog/on-policy-distillation/> — On-Policy Distillation \[\*\*]

  - \[meta]:

    - date: 10/28/2025
    - tags: \["distillation","rl"]
    - highlighted: true

- <https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/> — Evaluating frontier AI R\&D capabilities of language model agents against human experts

  - \[meta]:

    - date: 10/28/2025
    - tags: \["alignment"]

- <https://thenumb.at/Functions-are-Vectors/> — Functions are vectors \[\*\*]

  - \[meta]:

    - date: 10/27/2025
    - tags: \["programming","linear algebra"]
    - highlighted: true

- <https://assets.anthropic.com/m/71876fabef0f0ed4/original/reasoning_models_paper.pdf> — Reasoning models don’t always say what they think

  - \[meta]:

    - date: 10/09/2025
    - tags: \["interpretability"]

  - see also: <https://www.anthropic.com/research/reasoning-models-dont-say-think>

- <https://transformer-circuits.pub/2025/faithfulness-toy-model/index.html> — A Toy Model of Mechanistic (Un)Faithfulness

  - \[meta]:

    - date: 10/27/2025
    - tags: \["alignment"]

- <https://alignment.anthropic.com/2025/subliminal-learning/> — Subliminal Learning: Language models Transmit Behavioral Traits via Hidden Signals in Data

  - \[meta]:

    - date: 10/27/2025
    - tags: \["alignment"]

- <https://alignment.anthropic.com/2025/stress-testing-model-specs/> — Stress-testing model specs reveals character differences among language models

  - \[meta]:

    - date: 10/27/2025
    - tags: \["alignment"]

  - Models reveal their implicit value hierarchies in the value tradeoff scenarios we generate. By aggregating models’ decisions across our ‘high disagreement’ scenarios (the scenarios where responses vary the most across frontier models), we can identify clear patterns that distinguish different model families.

  - [Zhang et al. (2025)](#bib-zhang2025stress)

- <https://kyutai.org/next/codec-explainer> — Neural audio codecs: how to get audio into LLMs

  - \[meta]:

    - date: 10/26/2025
    - tags: \["audio","llm"]

- [youtube/v=1KRcs8XYUWo](https://www.youtube.com/watch?v=1KRcs8XYUWo) — Sequence-to-sequence models: [Connectionist](/thoughts/Connectionist-network) Temporal Classification

  - \[meta]:

    - date: 10/26/2025
    - tags: \["models","seq2seq"]

- <https://www.aleksagordic.com/blog/matmul> — Anatomy of high performance matmul kernels

  - \[meta]:

    - date: 10/24/2025
    - tags: \["gpu programming"]

  - <https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/>

- <https://maurice-weiler.gitlab.io/blog_post/cnn-book_1_equivariant_networks/> — Equivariant neural networks  –  what, why and how ?

  - \[meta]:

    - date: 10/24/2025
    - tags: \["neural network","equivariant"]

  - see also: [CNN](/thoughts/university/twenty-four-twenty-five/sfwr-4ml3/Convolutional-Neural-Network)

- <https://jameschen.io/jekyll/update/2024/02/12/mamba.html> — Mamba models

  - \[meta]:

    - date: 10/24/2025
    - tags: \["state space models","s4"]

  - see also: [state-space models](/thoughts/state-space-models#mamba)

- <https://www.tilderesearch.com/vignettes/gram-space> — Gram-Space Manifold Muon

  - \[meta]:

    - date: 10/24/2025
    - tags: \["optimizer"]

  - see also: [muon](/thoughts/muon)

- <https://x.com/jkminder/status/1980290860261732560> — Finetuning on narrow domains leaves traces behind. So can interpretability agents

  - \[meta]:

    - date: 10/21/2025
    - tags: \["interpretability"]

- <https://transformer-circuits.pub/2025/attention-qk/index.html> — Tracing Attention Computation Through Feature Interactions

  - \[meta]:

    - date: 10/09/2025
    - tags: \["interpretability"]

  - see also: [mechanistic interpretability](/thoughts/mechanistic-interpretability#qk-attributions)

- <https://www.goodfire.ai/research/replicating-circuit-tracing-for-a-simple-mechanism> — Replicating Circuit Tracing for a Simple Known Mechanism

  - \[meta]:

    - date: 10/09/2025
    - tags: \["interpretability"]

- <https://www.goodfire.ai/research/model-diff-amplification> — Discovering Undesired Rare Behaviors via Model Diff Amplification \[\*\*]

  - \[meta]:

    - date: 10/09/2025
    - tags: \["model diff","interpretability"]
    - highlighted: true

- <https://x.com/thesubhashk/status/1887138694546788556> — Helix representation in LLMs for additions capabilities

  - \[meta]:

    - date: 10/09/2025
    - tags: \["math","llm"]

- <https://www.goodfire.ai/blog/on-optimism-for-interpretability> — On Optimism for interpretability

  - \[meta]:

    - date: 10/09/2025
    - tags: \["interpretability"]

- <https://leimao.github.io/blog/CuTe-Tilers/> — CuTe megathread \[\*\*]

  - \[meta]:

    - date: 10/07/2025
    - tags: \["ml","compiler"]
    - highlighted: true

  - <https://leimao.github.io/article/CuTe-Layout-Algebra/>

  - <https://leimao.github.io/blog/CuTe-Inverse-Layout/>

  - <https://leimao.github.io/blog/CuTe-Blocked-Raked-Products/>

  - <https://leimao.github.io/blog/CuTe-Index-To-Coordinate/>

  - <https://leimao.github.io/blog/CUDA-Driver-Runtime-Load-Run-Kernel/>

- <https://hanlab.mit.edu/blog/svdquant-nvfp4> — SVDQuant Meets NVFP4: 4× Smaller and 3× Faster FLUX with 16-bit Quality on NVIDIA Blackwell GPUs

  - \[meta]:

    - date: 10/06/2025
    - tags: \["inference","optimization"]

- [youtube/v=i6Y2EelEC04](https://www.youtube.com/watch?v=i6Y2EelEC04) — iris: First-Class Multi-GPU Programming Experience in Triton \[\*\*]

  - \[meta]:

    - date: 10/06/2025
    - tags: \["compiler","gpu"]
    - highlighted: true

  - [GPU programming](/thoughts/GPU-programming#amd)

- <https://thinkingmachines.ai/blog/lora/> — LoRA Without Regret

  - \[meta]:

    - date: 10/06/2025
    - tags: \["ml","training","lora"]

- <https://hanlab.mit.edu/blog/streamingllm> — [Attention](/thoughts/Attention) sink keeps language models stable \[\*\*]

  - \[meta]:

    - date: 10/06/2025
    - tags: \["ml"]
    - highlighted: true

  - see also: [KV compression](/thoughts/KV-compression#streaming-llm)

  - Diagraph sink

  - OpenAI: attention\_probs=softmax(\[sink\_scalar,a1​,a2​,…,at​])

  - [Barbero et al. (2025)](#bib-barbero2025llmsattendtoken) have shown that attention sinks serve as “pressure valves” preventing what researchers call “over-mixing”—a pathological state where deep models processing long sequences blur important distinctions between tokens.

- <https://www.neuronpedia.org/graph/info> — The Circuits Research Landscape: Results and Perspective, Aug 2025 \[\*\*]

  - \[meta]:

    - date: 10/06/2025
    - tags: \["interpretability"]
    - highlighted: true

  - [mechanistic interpretability](/thoughts/mechanistic-interpretability)

- [lesswrong/llms-can-learn-about-themselves-by-introspection![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/L3aYFT4RDJYHbbsup/llms-can-learn-about-themselves-by-introspection) — LLMs can learn about themselves by introspection

  - \[meta]:

    - date: 10/05/2025
    - tags: \["alignment","ai"]

  - related to [Alignment](/thoughts/Alignment)

  - TLDR: We find that [LLMs](/thoughts/LLMs) are capable of introspection on simple tasks. We discuss potential implications of introspection for interpretability and the moral status of AIs.

    - I think this is largely [emergent behaviour](/thoughts/emergent-behaviour) based on [observer-expectancy effect](/thoughts/observer-expectancy-effect) through learnt patterns in RL/post-training paradigm.

    - LLMs can acquire knowledge that cannot be inferred from their training data. This challenges the view that LLMs simply imitate their training distributions. Instead, it appears that some LLMs have “privileged access” to certain facts about themselves and can use it to answer questions.

      - I wonder if the models grok based on what they understand OOD? We certainly don’t have a strong hypothesis on why model groks overall.

- [lesswrong/an-introduction-to-representation-engineering-an-activation![LessWrong](/static/favicons/lesswrong.avif)](https://www.lesswrong.com/posts/3ghj8EuKzwD3MQR5G/an-introduction-to-representation-engineering-an-activation) — Representation Engineering, an activation-based paradigm for controlling [LLMs](/thoughts/LLMs)

  - \[meta]:

    - date: 10/05/2025
    - tags: \["interpretability"]

- <https://www.thought-anchors.com/> — Thought Anchors in LLM Reasoning Traces

  - \[meta]:

    - date: 10/05/2025
    - tags: \["interpretability","reasoning","grpo"]

  - [2506.19143![arXiv](/static/favicons/arxiv.avif)](https://arxiv.org/abs/2506.19143) ([Bogdan et al., 2025](#bib-bogdan2025thoughtanchorsllmreasoning))

- <https://kipp.ly/transformer-inference-arithmetic/> — Transformers Inference Arithmetics \[\*\*]

  - \[meta]:

    - date: 10/05/2025
    - tags: \["napkin","inference"]
    - highlighted: true

  - see also [Transformers](/thoughts/Transformers#inference), [LLMs](/thoughts/LLMs)

- [youtube/v=kLiwvnr4L80](https://www.youtube.com/watch?v=kLiwvnr4L80\&t=868s) — Trends in Deep Learning, Bill Dally (NVIDIA)

  - \[meta]:

    - date: 10/05/2025
    - tags: \["argumentative","trend"]

  - [GPU programming](/thoughts/GPU-programming), [NVIDIA architecture notes](/lectures/420)

- <https://x.com/thesephist/status/1895887696268288119> — AI-centric interface design

  - \[meta]:

    - date: 10/03/2025
    - tags: \["design","interface"]

- <https://x.com/ZayneSprague/status/1836784332704215519> — To CoT or not to CoT

  - \[meta]:

    - date: 10/03/2025
    - tags: \["reasoning","interpretability"]

- <https://x.com/simonw/status/1840438066974228912> — NotebookLLM system prompt

  - \[meta]:

    - date: 10/03/2025
    - tags: \["prompting"]

- <https://x.com/karpathy/status/1841536804073439268> — Karpathy’s at GPU MODE’s IRL talk

  - \[meta]:

    - date: 10/03/2025
    - tags: \["systems","vllm"]

- <https://x.com/stephen_wolfram/status/1826692234554875979> — Deep-dive into ML

  - \[meta]:

    - date: 10/03/2025
    - tags: \["internal","ml"]

- <https://x.com/sarahookr/status/1834294208821428571> — inference-time not capture governance guardrails

  - \[meta]:

    - date: 10/03/2025
    - tags: \["alignment","policy","safety"]

- <https://x.com/banburismus_/status/1819354340290658725> — Tom McGrath’s questions about cross-layer superposition

  - \[meta]:

    - date: 10/03/2025
    - tags: \["interpretability"]

- <https://x.com/NeelNanda5/status/1850656772002120009> — Neel’s take on Anthropic’s crosscoders

  - \[meta]:

    - date: 10/03/2025
    - tags: \["interpretability","argumentative"]

- <https://x.com/AIatMeta/status/1851327605716435011> — layer skip in self-speculative decoding

  - \[meta]:

    - date: 10/03/2025
    - tags: \["inference"]

- <https://x.com/JungleSilicon/status/1866352582349750555> — embedding visualisation from Midjourney

  - \[meta]:

    - date: 10/03/2025
    - tags: \["latent space"]

- <https://x.com/sleenyre/status/1851519830375207309> — sae for flux-lens for exploring image embeddings

  - \[meta]:

    - date: 10/03/2025
    - tags: \["interpretability"]

- <https://x.com/jxmnop/status/1851706815244902691> — contextual document embeddings OSS

  - \[meta]:

    - date: 10/03/2025
    - tags: \["embeddings","latent space"]

- <https://x.com/JustinLin610/status/1861847752835248381> — QwQ reasoning models outperform o1

  - \[meta]:

    - date: 10/03/2025
    - tags: \["reasoning","rl"]

- <https://x.com/ch402/status/1874990808539275687> — Chris Olah on state of AI research

  - \[meta]:

    - date: 10/03/2025
    - tags: \["research"]

- <https://x.com/giansegato/status/1875944887973183785> — The opportunity is now, don’t believe in both extreme wrt to AI

  - \[meta]:

    - date: 10/03/2025
    - tags: \["ai","longtermism"]

- <https://x.com/nrehiew_/status/1876091138366652438> — ML with shape suffixes stylistic choice

  - \[meta]:

    - date: 10/03/2025
    - tags: \["ml","interpretability"]

- <https://x.com/behrouz_ali/status/1878859086227255347> — Titan, scaling Neural Memory

  - \[meta]:

    - date: 10/03/2025
    - tags: \["attention","architecture"]

- <https://x.com/vllm_project/status/1879979185474859303> — By yours truly

  - \[meta]:

    - date: 10/03/2025
    - tags: \["structured outputs","inference"]

- <https://x.com/flowersslop/status/1882241958397067677> — R1 having existential crisis

  - \[meta]:

    - date: 10/03/2025
    - tags: \["rl","reasoning","cot"]

- <https://x.com/rupspace/status/1877882538859078084> — Highway network

  - \[meta]:

    - date: 10/03/2025
    - tags: \["model architecture"]

  - see also: <https://rupeshks.cc/blog/skip.html>

- <https://rupeshks.cc/blog/skip.html> — Weighted Skip Connections are Not Harmful for Deep Nets

  - \[meta]:

    - date: 12/11/2025
    - tags: \["interpretability","scaling"]

- <https://x.com/VictorTaelin/status/1897108466243641399> — Claude Code optimize HVM3 to 328 MIPs per M4 Core

  - \[meta]:

    - date: 10/03/2025
    - tags: \["engineering","optimization"]

  - <https://gist.github.com/VictorTaelin/4f55a8a07be9bd9f6d828227675fa9ac>

- <https://x.com/karpathy/status/1937902205765607626> — Karpathy on “context engineering” over prompt engineering

  - \[meta]:

    - date: 10/03/2025
    - tags: \["context engineering"]

- <https://x.com/jkminder/status/1939790920326541601> — Model diffing on Chat vs. Base Model

  - \[meta]:

    - date: 10/03/2025
    - tags: \["interpretability","techniques"]

- <https://x.com/leloykun/status/1941067659157913625> — Adam with Aggressive Gradient Value/Norm Clipping ≈ Smoothed SignSGD/NormSGD

  - \[meta]:

    - date: 10/03/2025
    - tags: \["optimizer"]

- <https://x.com/Kimi_Moonshot/status/1944589115510734931> — Kimi K2 rough architecture

  - \[meta]:

    - date: 10/03/2025
    - tags: \["model architecture"]

- <https://x.com/jobergum/status/1945036230799892726> — ColBERT WASM for embeddings

  - \[meta]:

    - date: 10/03/2025
    - tags: \["engineering","bitter lesson"]

- <https://x.com/geoffreylitt/status/1950601513870499953> — Geoffrey Litt on capabilities debates

  - \[meta]:

    - date: 10/03/2025
    - tags: \["capabilities","hci"]

- <https://x.com/JayaGup10/status/1952871186888843528> — Integration landscape from labs

  - \[meta]:

    - date: 10/03/2025
    - tags: \["providers","inference"]

- <https://x.com/HSVSphere/status/1955714317816316150> — modern-infrastructure + AI

  - \[meta]:

    - date: 10/03/2025
    - tags: \["infrastructure"]

- <https://x.com/ShiqianMa/status/1971979845170315669> — Manifold Muon, as spectral GD \[\*\*]

  - \[meta]:

    - date: 10/03/2025
    - tags: \["optimizer"]
    - highlighted: true

- <https://x.com/GoodfireAI/status/1953903581075288470> — gpt-oss interpretability speed run at Goodfire

  - \[meta]:

    - date: 10/03/2025
    - tags: \["interpretability","models"]

  - experts actually seem to specialize - e.g. a “business expert” that activates most on business strategy & management topics.

  - just as Claude fixates on spiritual bliss after talking to itself for many turns, gpt-oss has its own attractor states: nonsense code & creative writing!

  - memoized during training: <https://x.com/jack_merullo_/status/1953860284638278043>

  - SAEs and features activated on mentions of LLaMA models

  - gpt-oss’ inability to act as a naive text completion model (vs. reverting to a chat format), even with jailbreaks, correlates with how subjectively cooked/slop-ish it feels

  - found that you can sometimes get interpolated reasoning levels from gpt-oss between “low”, “medium”, and “high” (but “none”, “ultra”, “infinite” don’t work)

  - basic contrastive steering of gpt-oss-20b, following Anthropic’s “persona vectors”: <https://x.com/MarkMBissell/status/1952919910134497332>

- <https://x.com/Zai_org/status/1954750596634054965> — GLM 4.5 Technical report

  - \[meta]:

    - date: 10/03/2025
    - tags: \["model architecture"]

  - Agentic workflow RL scale

  - <https://z.ai/blog/glm-4.5>

- <https://x.com/djcows/status/1955435075136606449> — Read your weights

  - \[meta]:

    - date: 10/03/2025
    - tags: \["interpretability"]

- <https://x.com/Kimi_Moonshot/status/1944589115510734931> — Kimi K2 architecture

  - \[meta]:

    - date: 10/03/2025
    - tags: \["model architecture"]

  - ![](./thoughts/images/kimi-architecture.webp)

- <https://x.com/nic__carter/status/1797635177973158182> — The neck-breaking speed of AI

  - \[meta]:

    - date: 10/03/2025
    - tags: \["development"]

- <https://blog.ezyang.com/2025/08/state-of-torch-compile-august-2025> — The state of torch.compile for training (Aug 2025)

  - \[meta]:

    - date: 12/03/2025
    - tags: \["compilers","scaling"]
    - later: true

- <https://blog.ezyang.com/2025/08/the-parallelism-mesh-zoo/> — The Parallelism Mesh Zoo

  - \[meta]:

    - date: 10/06/2025
    - tags: \["distributed","scaling"]

- <https://blog.ezyang.com/2025/12/code-review-as-human-alignment-in-the-era-of-llms> — Code review as human alignment, in the era of LLMs

  - \[meta]:

    - date: 01/08/2026
    - tags: \["engineering","productivity"]

- <https://x.com/JingyuanLiu123/status/1959093411283443726> — TPU vs GPU parallelism strategies \[\*\*]

  - \[meta]:

    - date: 09/26/2025
    - tags: \["hardware design"]
    - highlighted: true

- <https://x.com/GoodfireAI/status/1960378734852046859> — Adversarial examples affects feature share directions.

  - \[meta]:

    - date: 09/15/2025
    - tags: \["adversarial","interpretability","research"]

  - <https://x.com/livgorton/status/1960378437102657654>

  - [2508.17456![arXiv](/static/favicons/arxiv.avif)](https://arxiv.org/abs/2508.17456) ([Gorton & Lewis, 2025](#bib-gorton2025adversarialexamplesbugssuperposition))

- <https://x.com/keenanisalive/status/1964434335911858552> — autoencoder representations

  - \[meta]:

    - date: 10/03/2025
    - tags: \["visualisation"]

  - Geometric representation of encoders: maps a high-dimensional data x to low-dimensional latent z, then the decoder tries to map z back to x.

  - We _always_ learn a k-dimensional submanifold M ![](./thoughts/images/submanifold-mapping.webp)

  - See also: [diagrams](/thoughts/autoencoder-diagrams-intuition)

- <https://x.com/GoodfireAI/status/1965189414491168785> — SAE scaling law dynamics

  - \[meta]:

    - date: 10/03/2025
    - tags: \["interpretability","scaling law"]

  - scaling on feature manifolds

  - in terms of manifold discovered

- <https://x.com/_xjdr/status/1966215415027347856> — Qwen3-Next architecture difference with hybrid and Gated Delta Rule

  - \[meta]:

    - date: 10/03/2025
    - tags: \["model architecture"]

  - [2310.07707![arXiv](/static/favicons/arxiv.avif)](https://arxiv.org/abs/2310.07707) ([Devvrit et al., 2023](#bib-devvrit2024matformernestedtransformerelastic))

- <https://x.com/Grad62304977/status/1967548295816819184> — RL resources \[\*\*]

  - \[meta]:

    - date: 10/03/2025
    - tags: \["rl"]
    - highlighted: true

- <https://x.com/repligate/status/1968093240646889820> — Yud’s AI book

  - \[meta]:

    - date: 10/03/2025
    - tags: \["longtermism"]

- <https://x.com/mrsiipa/status/1968284758661894436> — The NVIDIA regime

  - \[meta]:

    - date: 10/03/2025
    - tags: \["gpu programming"]

  - I’m tired

- <https://x.com/jackminong/status/1968518159826305438> — PyTorch weirdness

  - \[meta]:

    - date: 10/03/2025
    - tags: \["ml framework"]

  - See also [Weight tying](/thoughts/Weight-tying)

- <https://x.com/josh_bickett/status/1725556267014595032> — What is an AI agent \[\*\*]

  - \[meta]:

    - date: 10/03/2025
    - tags: \["ai"]
    - highlighted: true

- <https://x.com/jkminder/status/1969082859311841413> — [sparse crosscoders](/thoughts/sparse-crosscoders) and non-linear representation dilemma

  - \[meta]:

    - date: 10/03/2025
    - tags: \["interpretability","hypothesis"]

- <https://x.com/karpathy/status/1973435013875314729> — Karpathy’s bitter lesson on Dwarkesh’s pod with Sutton

  - \[meta]:

    - date: 10/03/2025
    - tags: \["bitter lesson","retrospective"]

  - Stated plainly, today’s frontier LLM research is not about building animals. It is about summoning ghosts. You can think of ghosts as a fundamentally different kind of point in the space of possible intelligences

  - They are these imperfect replicas, a kind of statistical distillation of humanity’s documents with some sprinkle on top. They are not platonically bitter lesson pilled, but they are perhaps “practically” bitter lesson pilled, at least compared to a lot of what came before.

  - see also: <https://chatgpt.com/share/68dd6833-67c4-8007-8f37-331eb5bd9ee0>

- <https://x.com/deepcohen/status/1973191790602887544> — Central flows

  - \[meta]:

    - date: 10/03/2025
    - tags: \["optimizer"]

  - <https://centralflows.github.io/part1/>

    - How [gradient descent](/thoughts/gradient-descent) works?

- <https://x.com/nrehiew_/status/1973404310127124790> — DSA comparing to attention sink in long-context regime

  - \[meta]:

    - date: 10/03/2025
    - tags: \["attention","scaling law"]

  - see also: <https://x.com/nathancgy4/status/1973420757196873885>

  - [optimization](/thoughts/optimization#softmax) over 2048 tokens, thus QK weights magnitudes are preserved, and no “attention budget” are being given to useless tokens.

- <https://www.goodfire.ai/blog/painting-with-concepts> — Painting with concepts, via SAE for SDXL-turbo

  - \[meta]:

    - date: 05/30/2025
    - tags: \["interpretability","diffusion"]

- <https://www.goodfire.ai/blog/under-the-hood-of-a-reasoning-model> — Under the hood of a reasoning models, SAEs.

  - \[meta]:

    - date: 05/30/2025
    - tags: \["interpretability"]

- <https://wattenberger.com/thoughts/yay-embeddings-math> — creative with embeddings in writing

  - \[meta]:

    - date: 10/03/2025
    - tags: \["latent space"]

- <https://basilhalperin.com/essays/agi-vs-emh.html> — AGI timeline, Basil Halperin

  - \[meta]:

    - date: 05/30/2025
    - tags: \["agi","longtermism"]

- <https://optimists.ai/2024/03/10/deconstructing-bostroms-argument-for-ai-doom/> — Bostrom’s Argument for AI Doom

  - \[meta]:

    - date: 05/30/2025
    - tags: \["ai safety","longtermism"]

- <https://www.joelsimon.net/lluminate> — Creative exploration with LLM

  - \[meta]:

    - date: 05/30/2025
    - tags: \["lluminate"]

- [docs.google.com/1dG\[...\]IZw](https://docs.google.com/presentation/d/1dGA1Jpppv9BciOOrc95ZYRinsj_ZYrI-IQOIifHGIZw/edit?slide=id.p#slide=id.p) — Transformers in Diffusion Models for Image Generations and Beyond

  - \[meta]:

    - date: 07/13/2025
    - tags: \["model architecture"]

- <https://discuss.vllm.ai/t/numerical-difference-between-vllm-logprobs-and-huggingface-logprobs/151> — Numerical difference between [vLLM](/thoughts/vllm) logprobs and HF logprobs

  - \[meta]:

    - date: 07/13/2025
    - tags: \["inference","numerical stability"]

- <https://jeremybernste.in/writing/deriving-muon> — Deriving [Muon](/thoughts/muon)

  - \[meta]:

    - date: 08/02/2025
    - tags: \["optmizer"]

  - [KellerJordan/modded-nanogpt](https://github.com/KellerJordan/modded-nanogpt)

  - <https://kellerjordan.github.io/posts/muon/>

- <https://www.neuronpedia.org/graph/info> — The circuit analysis research landscape

  - \[meta]:

    - date: 08/05/2025
    - tags: \["fruit","interpretability"]

  - [mechanistic interpretability](/thoughts/mechanistic-interpretability#attribution-graph)

  - On Biology of LLMs

  - Futures and directions of interpretability [research](https://www.neuronpedia.org/graph/info#section-directions-for-future-work)

    - Influence functions, or training data attributions methods
    - Scaling interpreting CoT, exempli gratia [Docent](https://transluce.org/introducing-docent)

- <https://ethanding.substack.com/p/ai-subscriptions-get-short-squeezed> — Tokens are getting more expensive

  - \[meta]:

    - date: 08/05/2025
    - tags: \["inference economics"]

- [docs.google.com/1NO\[...\]X80](https://docs.google.com/presentation/d/1NOrUVZNkcKHom5ih5uqqOSrV4Vi8KZRN8LYaAeADX80/edit?slide=id.g3724263dfb4_0_68#slide=id.g3724263dfb4_0_68) — Scaling MoE with llm-d and vLLM

  - \[meta]:

    - date: 08/05/2025
    - tags: \["inference","scaling"]

- <https://www.anthropic.com/research/persona-vectors> — Persona vector

  - \[meta]:

    - date: 09/18/2025
    - tags: \["interpretability"]

- <https://www.tilderesearch.com/blog/momoe> — MoMoE: Memory-optimized Mixture of Experts

  - \[meta]:

    - date: 10/03/2025
    - tags: \["momoe"]

  - See also: [triton-based](https://github.com/shawntan/scattermoe) implementation of Sparse MoE

  - Qwen3 modular fused: [woct0rdho/transformers-qwen3-moe-fused](https://github.com/woct0rdho/transformers-qwen3-moe-fused/blob/master/qwen3_moe_fused/modular_qwen3_moe_fused.py)

- [docs.google.com/1ZV\[...\]YtI](https://docs.google.com/document/d/1ZV73D2vgaj2yu_tjN3TVOP6QVLWVPXJB2rrqSZQxYtI/edit?usp=drivesdk) — AI research overview

  - \[meta]:

    - date: 08/14/2025
    - tags: \["edit"]

- <https://timdettmers.com/2022/08/17/llm-int8-and-emergent-features/> — [LLMs](/thoughts/LLMs).int8() and Emergent Features

  - \[meta]:

    - date: 08/16/2025
    - tags: \["properties","inference"]

  - [2110.02861![arXiv](/static/favicons/arxiv.avif)](https://arxiv.org/abs/2110.02861) ([Dettmers et al., 2022](#bib-dettmers20228bitoptimizersblockwisequantization))

- <https://www.math.uwaterloo.ca/~hwolkowi/matrixcookbook.pdf> — Matrix cookbook

  - \[meta]:

    - date: 08/16/2025
    - tags: \["books","matmul"]

- [xjdr-alt/llmri](https://github.com/xjdr-alt/llmri) — LLM varentropy versus. entropy plot

  - \[meta]:

    - date: 08/16/2025
    - tags: \["entropy","inference","sampler"]

- <https://jax-ml.github.io/scaling-book/> — JAX scaling book \[\*\*]

  - \[meta]:

    - date: 08/16/2025
    - tags: \["training","large scale"]
    - highlighted: true

- <https://jaxformer.com/> — Training LLMs with Jax

  - \[meta]:

    - date: 10/07/2025
    - tags: \["training","large scale"]

  - made by [Cohere](https://cohere.com/)

- <https://nanotron-ultrascale-playbook.static.hf.space/index.html> — The Ultra-Scale Playbook: Training LLMs on GPU Clusters \[\*\*]

  - \[meta]:

    - date: 10/04/2025
    - tags: \["books","scaling law"]
    - highlighted: true

- <https://www.jeremykun.com/2023/08/10/mlir-getting-started/> — [MLIR](/thoughts/MLIR) introduction \[\*\*]

  - \[meta]:

    - date: 08/16/2025
    - tags: \["compiler","ml"]
    - highlighted: true

- <https://huggingfacefw-blogpost-fineweb-v1.static.hf.space> — FineWeb: decanting the web for the finest text data at scale. \[\*\*]

  - \[meta]:

    - date: 08/16/2025
    - tags: \["datasets"]
    - highlighted: true

- <https://www.cs.toronto.edu/~duvenaud/distill_bayes_net/public/> — [Bayesian Neural Network](/thoughts/Bayesian-Neural-Network) \[\*\*]

  - \[meta]:

    - date: 08/16/2025
    - tags: \["ml"]
    - highlighted: true

- <https://www.lei.chat/posts/triton-linear-layout-concept/> — Triton Linear Layout

  - \[meta]:

    - date: 08/16/2025
    - tags: \["abstract algebra"]

- <https://research.colfax-intl.com/cutlass-tutorial-writing-gemm-kernels-using-tensor-memory-for-nvidia-blackwell-gpus/> — GEMM kernels on Blackwell GPUs

  - \[meta]:

    - date: 08/16/2025
    - tags: \["gpu programming"]

- [docs.google.com/1o9\[...\]mtY](https://docs.google.com/document/d/1o9ZZEFofxb-dJ_cDTi_c3riOJOm4higcSGpvKKNdmtY/edit?tab=t.0#heading=h.nc4nxxczgw4w) — Structural tags in xgrammar (to be used in [vLLM](/thoughts/vllm))

  - \[meta]:

    - date: 08/16/2025
    - tags: \["inference","structured outputs"]

- [drive.google.com/1l54BwUi07JnqwB5\_iHCVVZ3TnR05acDm](https://drive.google.com/file/u/1/d/1l54BwUi07JnqwB5_iHCVVZ3TnR05acDm/view?usp=sharing) — The Illusion of The Illusion of Thinking

  - \[meta]:

    - date: 08/16/2025
    - tags: \["rl","contrarian"]

  - see also [2506.06941![arXiv](/static/favicons/arxiv.avif)](https://arxiv.org/abs/2506.06941) ([Shojaee et al., 2025](#bib-shojaee2025illusionthinkingunderstandingstrengths))

- <https://huggingface.co/blog/rearchitecting-uploads-and-downloads> — Rearchitecting Hugging Face Uploads and Downloads

  - \[meta]:

    - date: 08/16/2025
    - tags: \["engineering"]

- [docs.google.com/1p-\[...\]32I](https://docs.google.com/document/d/1p-ggQV3vVWIQuCccXEl1fD0thJOgXimlbBpGk6FI32I/edit?tab=t.0#heading=h.y0ohi6l5z9qn) — MATS 9.0 Winter 2025

  - \[meta]:

    - date: 10/03/2025
    - tags: \["interpretability","research"]

  - <https://x.com/NeelNanda5/status/1950344397075456438>

- <https://nousresearch.com/measuring-thinking-efficiency-in-reasoning-models-the-missing-benchmark/> — Measuring reasoning model thinking efficiency

  - \[meta]:

    - date: 08/16/2025
    - tags: \["rl","efficiency","cot"]

- <https://www.seangoedecke.com/great-software-design/> — Great software design looks underwhelming

  - \[meta]:

    - date: 08/18/2025
    - tags: \["design","engineering"]

- <https://openai.com/index/deep-double-descent/> — Deep double descent \[\*\*]

  - \[meta]:

    - date: 08/21/2025
    - tags: \["optimizer","ml"]
    - highlighted: true

- <https://horace.io/brrr_intro.html> — Deep Learning from first principle

  - \[meta]:

    - date: 08/28/2025
    - tags: \["gpu programming","scaling law"]

- <https://colah.github.io/posts/2014-03-NN-Manifolds-Topology/> — Neural Network, and Manifolds \[\*\*]

  - \[meta]:

    - date: 08/28/2025
    - tags: \["manifolds","interpretability","topology"]
    - highlighted: true

- <https://substack.com/@jakeeaton/note/c-140561607> — AI skeptics unconsciously anthropomorphize LLMs in their critiques, like their anger is directed more at a naive zoomer intern who can’t infer the context of your ask. \[—]

  - \[meta]:

    - date: 09/07/2025
    - tags: \["longtermism"]

  - it’s funny how many AI skeptics unconsciously anthropomorphize LLMs in their critiques, like their anger is directed more at a naive zoomer intern who can’t infer the context of your ask, rather than a computer program that can work small miracles if you have the patience to learn how it thinks.

    > no one gets mad at the limitations of their microwave, or even a stats package, the way they get worked up over a machine in a box that four years ago couldn’t do addition and today wins math olympiads. in that anger is an implicit assumption: that it should know better — and of something on the other end that’s far more than a machine, if not quite yet a soul

- <https://ghost.oxen.ai/why-grpo-is-important-and-how-it-works/> — Why GRPO is important and how it works \[\*\*]

  - \[meta]:

    - date: 09/10/2025
    - tags: \["rl","training"]
    - highlighted: true

  - pinned: true

- <https://mp.weixin.qq.com/s/h1cFYDNxcHC30APcarF47A> — [vLLM](/thoughts/vllm) from scratch

  - \[meta]:

    - date: 09/21/2025
    - tags: \["inference"]

  - 1.3: <https://mp.weixin.qq.com/s/BdWG6_ZTaGRknmsbGfFkMQ>

  - 1.2: <https://mp.weixin.qq.com/s/8BVEVPPqDQhQ2l8L90dMNQ>

- <https://ma-lab-berkeley.github.io/deep-representation-learning-book/> — Learning Deep Representations of Data Distributions

  - \[meta]:

    - date: 09/28/2025
    - tags: \["ml"]

  - see also: [Buchanan et al. (2025)](#bib-ldrdd2025)

- <https://www.julian.ac/blog/2025/09/27/failing-to-understand-the-exponential-again/> — Failing to Understand the Exponential, Again

  - \[meta]:

    - date: 09/29/2025
    - tags: \["bitter lesson","training"]

- <https://hazyresearch.stanford.edu/blog/2025-09-28-tp-llama-main> — We Bought the Whole GPU, So We’re Damn Well Going to Use the Whole GPU

  - \[meta]:

    - date: 10/03/2025
    - tags: \["inference"]

- <https://hazyresearch.stanford.edu/blog/2025-05-27-no-bubbles> — Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B \[\*\*]

  - \[meta]:

    - date: 11/12/2025
    - tags: \["kernel","inference"]
    - highlighted: true

