Jul 27Tuesday, July 28, 2026 · all days
1.Our position on open-weights models(anthropic.com)
943 points by surprisetalk 13 hours ago | 1362 comments | permalink
tl;dr: Anthropic CEO Dario Amodei clarifies that the company does not support banning open-weights models, including Chinese ones, and calls such bans ineffective for national security. Instead, he advocates for three measures: restricting powerful chip exports to China, cracking down on industrial-scale model distillation, and mandating pre-release safety testing for all sufficiently capable models regardless of origin or whether they're open or closed.
HN Discussion:
  • Anthropic's stance is self-serving protectionism disguised as safety concern to eliminate competition
  • Mandatory safety testing is effectively a backdoor ban on open-weights models via gatekeeping
  • Anthropic's moral posturing is hypocritical given its ties to military and current US regime
  • The article's logic is internally inconsistent, particularly on bans working for chips but not models
  • Open-weight proliferation poses genuine bioweapon/cyber risks that critics ignore
2.Benchmarking Opus 5 on SlopCodeBench(github.com)
306 points by dhorthy 12 hours ago | 70 comments | permalink
tl;dr: SlopCodeBench is a long-horizon coding benchmark where models must evolve a codebase across sequential checkpoints without seeing future requirements upfront, making it a good proxy for real maintenance work. On a 17-checkpoint subset, Opus 5 got 24% strict pass (vs 6% for Opus 4.8 and Sonnet 5), but no model completed any full challenge cleanly, and all showed rising complexity, duplication, and verbosity over time. The author argues this finally provides hard data for the intuition that current frontier models can't be trusted to run "lights-off" on iterative software work without human steering.
HN Discussion:
  • Benchmark uniquely captures long-horizon maintenance work and code cleanliness better than alternatives
  • Labs should use this in RL pipelines to prioritize reducing code complexity
  • ~Results need human baseline comparison to be properly interpreted
  • Personal experience confirms Opus 5 is only a modest improvement over predecessors
  • Suggestions to extend methodology (test ordering, PR review integration, adversarial prompting)
3.Using an open model feels surprisingly good(matthewsaltz.com)
288 points by msaltz 8 hours ago | 111 comments | permalink
tl;dr: A Modal employee set up opencode with a self-hosted Kimi K3 inference endpoint on Modal instead of upgrading their Claude plan, and found the experience unexpectedly liberating. They describe owning the endpoint and controlling their own data as feeling "freeing"—akin to switching from a fancy IDE back to vim—despite not previously being an open-source enthusiast.
HN Discussion:
  • The post is thinly veiled self-promotion or an advertisement for Modal
  • Frontier models like Claude are better for vague, one-shot prompts, limiting open models' appeal
  • Open/smaller models feel surprisingly capable and enable a better flow state, echoing the vim analogy
  • Freedom from vendor lock-in and controlling your own data is the real value
  • ~Cost comparison data is missing and would make the case more compelling
4.Watching Go's new garbage collector move through the heap(theconsensus.dev)
233 points by matheusmoreira 3 days ago | 29 comments | permalink
tl;dr: Summary not available
HN Discussion:
  • ~Article was interesting but ended abruptly, leaving readers wanting more
  • Heap visualization effectively demystifies GC behavior and highlights memory layout importance
  • Appreciation for the manual object-copying optimization technique described
  • ~Would benefit from practical guidance on observing GC in one's own applications
  • Questions the efficiency benefit if objects already reside on the same page
5.Paged Out #9 [pdf](pagedout.institute)
248 points by laurensr 20 hours ago | 27 comments | permalink
tl;dr: Summary not available
HN Discussion:
  • Highlights specific articles/pages as enjoyable or funny reads
  • Praises the zine's design and hacker-culture aesthetic, comparing to classics like 2600 and Phrack
  • Adds technical context noting a piece rediscovers Wang's 1960s work on computable tilings
  • Interested in purchasing print editions and asking about availability
  • General fandom expression for the zine
6.Netflix employee fired for sharing personal details in retreat trust exercise(nypost.com)
348 points by softwaredoug 11 hours ago | 340 comments | permalink
tl;dr: A former Netflix VP, Kevin Baillie, is suing the company after being fired for disclosing during a corporate "Vulnerability-Trust exercise" that he had undergone medically supervised ketamine therapy for depression following his mother's death. Netflix's attorney reportedly confirmed the ketamine disclosure factored into his termination, which also cited profanity and drinking concerns—despite Baillie's claim that the Eyeline Studios CEO cultivated an alcohol-heavy work culture. Baillie is seeking damages and a jury trial, alleging he was denied up to a year of severance.
HN Discussion:
  • Corporate trust exercises are manipulative traps designed to exploit vulnerable employees
  • Offsite retreats and forced drinking cultures are inherently problematic and dangerous
  • Never disclose personal information to employers because they aren't your friends
  • Netflix will likely lose or settle this lawsuit due to weak justification
  • Questioning whether ketamine use could legitimately impact his specific job role
7.A missing underscore sent innocent man to prison for 18 months(arstechnica.com)
318 points by quantified 12 hours ago | 182 comments | permalink
tl;dr: A Wisconsin child-luring investigation targeting Kik user "fus__ro_dah" (two underscores) went awry when police subpoenaed records for "fus_ro_dah" (one underscore), leading them to innocent Nova Scotia man Brandon Klayme. Despite no incriminating evidence on his devices, Klayme was convicted in 2023 and served 18 months in prison. The underscore error was discovered during appeal preparation, and the Nova Scotia Court of Appeal has now overturned his conviction, declaring him factually innocent.
HN Discussion:
  • Highlights the lack of evidence connecting Klayme to the crime, reinforcing the injustice
  • Questions whether compensation was provided for the wrongful conviction and lasting damage
  • Criticizes the defense attorneys for failing to challenge the prosecution's weak case
  • ~Alarmed that a mere username match could frame anyone, questioning if details are missing
  • ~Questions what evidence actually led to conviction, suggesting the article omits key details
8.Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped(techdirt.com)
316 points by cdrnsf 16 hours ago | 128 comments | permalink
tl;dr: A judge dismissed Google's DMCA 1201 lawsuit against SerpAPI for scraping search results, ruling that Google's "SearchGuard" anti-bot system doesn't qualify as a copyright protection measure since search results largely contain non-copyrighted public info, and Google lacks authorization from actual copyright holders to deploy it on their behalf. Google can refile narrower claims covering content it actually owns (like Knowledge Panel material), but the broader attempt to weaponize the DMCA against scraping failed—an ironic outcome given Google built its business on scraping the web.
HN Discussion:
  • Google's deprecation of its search API created demand for scrapers like SerpAPI
  • Google's lawsuit was a bad-faith bullying tactic against a smaller company
  • The ruling is ironic given Google built its empire by scraping the web
  • ~Whether search results qualify as copyrightable data is legally gray, not clear-cut
  • Google still wins by establishing precedent that protects its own scraping practices
9.MAI-Cyber-1-Flash inside MDASH(microsoft.ai)
233 points by migmartri 18 hours ago | 109 comments | permalink
tl;dr: Microsoft launched MAI-Cyber-1-Flash, a compact security-focused model integrated into its MDASH vulnerability remediation harness, designed to handle ~90% of security tasks while reserving larger models like GPT-5.4 for harder cases. The combined system scores 96% on CyberGym while cutting costs 50% versus their previous best MDASH configuration. Microsoft also introduced Perception, an agentic security system providing agent teams for continuous monitoring, patching, and threat response workflows.
HN Discussion:
  • Microsoft's data advantage may only apply to Microsoft's own ecosystem, not broader environments
  • Access and usability of the product is unclear or hidden behind corporate bureaucracy
  • ~Defense is fundamentally asymmetric; needs formal verification or immune-system-like monitoring
  • Skepticism about Microsoft's execution and product naming based on past track record
  • Requests for openness — open weights, open source models, training data, and benchmarks
10.How is the Bun rewrite in Rust going?(lockwood.dev)
472 points by tomlockwood 23 hours ago | 372 comments | permalink
tl;dr: Six weeks after Bun's much-hyped "Rust rewrite" was supposedly merged, there's still no release tag, open PRs from Claude have nearly doubled to 2,475, and Anthropic employees are increasingly involved in ongoing work. The author argues the widely-cited $165K rewrite cost is misleading—actual spending likely approaches $800K when factoring in continued Claude credits and CI/CD—and that the project is being used as marketing proof of AI coding capability to support Anthropic's valuation. Similar AI-rewrite projects (Anthropic's C compiler, Cursor's FastRender browser) have gone silent for months.
HN Discussion:
  • Rewrite is actually shipping and going well; article's premise is misleading or outdated
  • Post-rewrite slowdown is normal and expected, not evidence of failure
  • LLM rewrites create illusion of progress but real costs and problems come later
  • Original issues were self-inflicted, undermining justification for the rewrite
  • ~Discourse is driven by ideological priors about AI rather than facts
11.Kimi-K3 on HuggingFace(huggingface.co)
1342 points by nateb2022 1 day ago | 531 comments | permalink
tl;dr: Moonshot AI released Kimi K3, an open-weight 2.8T-parameter MoE model (104B activated) with native multimodality and a 1M-token context window, built on new Kimi Delta Attention (KDA) and Attention Residuals architecture with MXFP4 quantization-aware training. Benchmarks show it competitive with Claude Opus 4.8, GPT-5.6, and Claude Fable 5 across reasoning, coding, agentic, and vision tasks, positioning it as the first open 3T-class frontier model. Weights are available on HuggingFace under the Kimi K3 License, with vLLM/SGLang deployment support and an OpenAI/Anthropic-compatible API.
HN Discussion:
  • Hosting costs and pricing dynamics for a 3T model will be revealing but expensive
  • Consumer hardware is poorly shaped for running large models locally
  • The real value is customization, fine-tuning, and IP sovereignty from open weights
  • License restrictions and revenue thresholds limit true openness of the release
  • Model has identity confusion claiming to be Claude, casting doubt on originality
12.PGSimCity - How PostgreSQL Works(nikolays.github.io)
909 points by jonbaer 1 day ago | 89 comments | permalink
tl;dr: PGSimCity is an educational, SimCity-styled interactive visualization that models PostgreSQL's internal engine to help users understand how it works. It's an early, unreviewed prototype that may contain inaccuracies, and contributions via issues or pull requests are welcomed. The project is independent and not affiliated with Electronic Arts.
HN Discussion:
  • ~Concept is appreciated but visualization is too busy and confusing to be effective
  • ~Would benefit from being interactive, e.g. entering a query and walking through the flow
  • Excellent educational approach that could be applied to other complex technical domains
  • Skeptical of accuracy since it was vibe-coded quickly with AI assistance
  • AI-assisted learning tools like this are a great new way to explore complex topics
13.The computer that helped win World War II(spectrum.ieee.org)
203 points by baruchel 5 days ago | 78 comments | permalink
tl;dr: In 1941, British codebreakers intercepted a new German encrypted teletype system ("Tunny," made by Lorenz) that was more advanced than Enigma. After Bill Tutte deduced the machine's workings and devised a statistical decryption method, engineer Tommy Flowers built Colossus—the world's first large-scale programmable electronic digital computer, using ~2,000 vacuum tubes—to automate the process. Ten Colossi ran at Bletchley Park by war's end, decrypting high-level German communications and likely shortening the war by months. Colossus is now being recognized as an IEEE Milestone.
HN Discussion:
  • Recommends visiting Bletchley Park and the National Museum of Computing to see Colossus firsthand
  • Shares historical context about Colossus being destroyed post-war and later rebuilt from memory
  • Article ignores crucial Polish mathematicians' contributions to codebreaking
  • Colossus was a special-purpose calculator, not a true general-purpose computer
  • Clarifies that Colossus and the Bombe attacked different ciphers, explaining architectural differences
14.Removing React.js from the codebase and adapting Htmx for UI interactivity (2023)(misago-project.org)
245 points by Ralfp 1 day ago | 180 comments | permalink
tl;dr: Misago (Django forum software) is removing React.js in favor of HTMX to eliminate the duplication of rendering logic between Django templates and React components, which slowed development, complicated plugins, and duplicated translations. Since forum interactivity is largely isolated to specific UI elements, HTMX's declarative HTML-swapping approach fits better than a full SPA. Early migration steps (account settings, threads lists) have already cut misago.js by ~85kb gzipped, with a full transition planned through late 2024.
HN Discussion:
  • HTMX is a great fit for forum software and server-rendered content
  • HTMX works well across many web apps including PWAs and replaces React/Vue effectively
  • HTMX struggles with performance when responses become large and complex
  • HTMX lacks proper tooling like Storybook for frontend development
  • Alternative server-side rendering approaches like PyView or Django REST APIs are worth considering
15.Show HN: Physically accurate black hole you can put in your room(blackhole.plav.in)
470 points by aplavin 4 days ago | 181 comments | permalink
tl;dr: Summary not available
HN Discussion:
  • Enthusiastic praise for the visualization's quality and smoothness
  • Critique that the 'physically accurate' claim is misleading science communication
  • ~Technical questions about specific accuracy details like accretion disk brightness asymmetry
  • Reports of technical issues with AR mode on certain devices
  • Tangential reflections on VR experiences and phobias triggered by large objects
16.Should you wash your solar panels?(incoherency.co.uk)
234 points by surprisetalk 22 hours ago | 234 comments | permalink
tl;dr: Cleaning visibly dusty solar panels yielded a 2-5% power output increase (worth ~£60-150/year) by comparing power ratios between two banks before and after washing one. Upgrading the 15-year-old 3.7kW system to modern panels would boost output ~60% and pay back in 3 years, but doing so would forfeit a lucrative legacy feed-in tariff, making the upgrade economically unattractive.
HN Discussion:
  • Cleaning yields much larger gains (10-40%) than the article's modest 2-5% figure
  • ~Long-term owners report minimal degradation and rely on rain, questioning need for washing
  • Technical explanations for the measurement curve shape (panel matching, cooling effects)
  • Economics of cleaning are oversold; article correctly highlights modest real returns
  • ~Creative workarounds like separate new systems to preserve feed-in tariff
17.French firefighters face 'pyrocumulonimbus' for first time(france24.com)
453 points by saaaaaam 1 day ago | 360 comments | permalink
tl;dr: French firefighters are battling a pyrocumulonimbus ("fire cloud") for the first time in France—a phenomenon previously confined to Australia and North America, where extreme ground heat drives a convective column that generates its own winds, lightning, and unpredictable multi-directional fire fronts. Officials say the blaze cannot be fought directly and have declared "operational impossibility," resorting to defensive tactics and hoping either for prolonged heavy rain or to steer the fire toward the sea.
HN Discussion:
  • Article's 'first time' claim is inaccurate per French sources citing prior events
  • Terminology quibble: should be pyrocumulus not pyrocumulonimbus since fire clouds don't rain
  • Background context on why the Landes forest is uniquely vulnerable to fire
  • Firsthand accounts and comparisons confirming the fire's severity and phenomenon
  • Broader climate crisis reflections, calling for geoengineering or noting widespread fires
18.What is happening to jobs? Separating AI hype from reality(siepr.stanford.edu)
297 points by pod_krad 2 days ago | 376 comments | permalink
tl;dr: Aggregate employment data shows little evidence of an AI-driven jobs apocalypse, though entry-level white-collar workers—especially in software and customer service—appear to be experiencing weaker hiring since 2022, with causes hard to disentangle from post-pandemic hiring corrections and interest rate hikes. Experimental studies show AI generally boosts worker productivity (particularly for less-skilled workers), but gains are uneven and haven't shown up in aggregate statistics. Firm adoption is accelerating but concentrated in tech and information-heavy sectors, with most companies still in pilot phases—suggesting AI's labor market impact, while real in pockets, remains early and uncertain.
HN Discussion:
  • Study data is outdated because coding/general agents only became effective in late 2024/2025
  • ~AI productivity gains are concentrated among top performers, amplifying existing inequality
  • AI helps less experienced workers but hinders experienced ones, matching article's claim
  • Organizational inertia masks real AI impacts that are unevenly distributed
  • ~Job market and AI hype from executives are disconnected from actual fundamentals
19.Scriptc by Vercel: TypeScript-to-Native compiler, no JavaScript engine in binary(github.com)
275 points by maxloh 1 day ago | 153 comments | permalink
tl;dr: Scriptc is a Vercel tool that compiles standard TypeScript into small native binaries (170KB–3MB) with ~2ms startup, no bundled JS engine by default, using clang and a custom C runtime. It uses the real TypeScript compiler for type-checking and covers most of the language plus Node APIs (fs, http, tls, fetch, etc.); code that can't be statically compiled is either rejected with diagnostics or run via an embedded QuickJS engine in an opt-in `--dynamic` mode. Correctness is enforced via differential testing against Node (byte-for-byte stdout/stderr matching) and AddressSanitizer runs.
HN Discussion:
  • Vercel is chasing hype with a vibecoded, unmaintained project that solves no real problem
  • TypeScript's value comes from the npm ecosystem, making native compilation impractical for real projects
  • Skepticism about progress speed compared to similar serious projects like Porffor
  • ~Benchmarks confirm tradeoffs: slower runtime but much better startup, memory, and binary size
  • Promising approach that enables TypeScript-to-native workflows and potential mobile/embedding use cases
20.Decathlon Germany adds Wero payment option to decathlon.de website(sgieurope.com)
328 points by doener 18 hours ago | 217 comments | permalink
tl;dr: Decathlon Germany enabled Wero, the European Payments Initiative's account-to-account payment scheme, on decathlon.de as of July 20, becoming the first market in Decathlon's international footprint to adopt it and joining merchants like Eventim Live ahead of Lidl and Rossmann. The move lets Decathlon bypass the 1-2% fees charged by Visa/Mastercard while tying payments to its loyalty program, with in-store rollout across ~110 German locations planned once Wero's POS functionality arrives in 2026-2027. EPI is funding a €10 voucher promotion in October to drive adoption, as Wero (~56M users) still struggles with the classic two-sided network problem.
HN Discussion:
  • Wero is a valuable European payment system building on solid SEPA foundations
  • Wero fails at true independence by requiring iOS/Android with Google Play services
  • Firsthand Wero checkout experience was smooth and impressive
  • Wero is just marketing over bank transfers with higher fees than free alternatives like EPC QR
  • ~Other payment systems (Blik, WeChat/Alipay, iDEAL) already do this better or first