Tags / Model research
LLM
- Failure’s Worth, Through V2er Metaphors: Waterfall, Helicopter, Map A V2EX thread on Fields medalists warning about AI cracking math problems. The real gold is the metaphors—hiking vs helicopter, copying exam answers, fog-of-war maps, and inventing the crane.
- Distillation Through V2er Metaphors: Fishbowls, Kitchens, and Copying A V2EX thread asked if model distillation is like fishing from someone else’s bucket. The replies turned an abstract term into vivid metaphors—and a clean takeaway.
- When Average People Feel AI: Electricity for Knowledge Work, a Rounding Error Outside Nathan Lambert: the AI revolution may last a century, but daily life barely notices. Benefits are too indirect; politics looks like Engels’ pause. Deep adoption, not chat, marks the real shift.
- An Alien Mind: Alignment Unsolved—Don’t Scale Flat-Out OpenAI’s chief scientist: AI is a grown alien intellect; goal ≠ value alignment; CoT monitoring is weakening. No lab has solved alignment enough for max-speed scaling.
- ChatGPT Improves Answers; Critical Thinking Widens Ideas A 1,000-student RCT from Bocconi and OpenAI: AI raises polish; causal-reasoning training raises originality. Old rubrics miss half the signal.
- Will AI Shrink Total Jobs? Demand vs Productivity A V2EX thread argues jobs track demand, not productivity — and human desire forever outruns AI. The comments tear into effective demand, prices and deflation.
- What Self-Improving Agents Actually Improve: Warp vs OpenAI Warp’s two-skill feedback loop is shippable; OpenAI’s shared-memory incident is not proven RSI. Skills compound from human labels—not weight updates.
- DeepSeek Engram: conditional memory as a new sparsity axis DeepSeek Engram adds O(1) lookup for the conditional memory MoE never had. At equal parameter and compute budget, reasoning gains more than rote knowledge.
- Software Is Not Files: Finding Materials ≠ Having Knowledge vivo Xiao Bo's KDC part 2: representation is not knowledge. A top RAG hit on a stale refund policy gave a confident wrong answer — governance beats tuning.
- Nvidia's Risky Business: AI Funding Enters the Danger Zone Stratechery deep dive: from 1873 railroad bonds to hyperscaler debt and Nvidia's $500B GPU financing platform, each funding layer is riskier than the last.
- The Defense Window Is Closing: Daybreak and GPT-5.6-Cyber OpenAI answers AI-scale attacks with Daybreak Blue/Red tiers and GPT-5.6-Cyber, taking advanced cyber task completion from 1.5% to 95% and finding real CVEs.
- Green Dashboards, Worse Answers: Latency Is Correctness Flat error rates hide thinner RAG context and truncated agent loops. Why TTFT, tail latency and retrieval metrics matter for answer quality, not just speed.
- Before You Let Agents Run Loose: AI Context Architecture Context architecture is not the RAG pipeline — it's guardrails, scopes, trust scores and human-in-the-loop. Stack Overflow on infrastructure vs architecture.
- Wide + Deep: why a 4B model can punch up DeepSeek taught models to think deep. A Tsinghua team says deep isn't enough — you also need wide. How a 4B setup can stand next to a 671B.