2026-07-28 Hacker News Top Stories #
- Kimi-K3 是 Moonshot AI 开源的多模态代理模型,总参数量 2.8T、激活参数 104B,支持文本图像输入和百万 token 上下文,引发关于其高昂部署成本与自建可行性的讨论。
- 美国公民因使用自动擦除数据的 GrapheneOS 手机在机场搜查时拒绝解锁被起诉,案件引发边境执法权力与宪法权利的激烈争论。
- PGSimCity 是一个用城市建筑隐喻 PostgreSQL 内部机制的交互式 3D 模拟器,帮助理解数据库性能行为,其可视化和 AI 辅助开发方式引发热议。
- AI 公司大规模购买并扫描后销毁稀有书籍,虽被裁定合理使用,却引发对版权保护、绝版作品及文化保存的深刻反思。
- Bun 开发者用 Rust 重写项目引发对真实成本与发布进度的质疑,文章揭露 CI/CD 持续 86 天未发布新版本,实际花费可能远超公布数字。
- 法国西南部野火首次产生极其危险的“火积云”现象,火势不可预测迫使消防员转入防御状态,引发对气候灾难加剧和人类应对能力的忧虑。
- Decker 是一个继承 HyperCard 精神的免费多媒体创作平台,内置脚本语言和 SQL 支持,可导出为 HTML,其限制激发创造力的理念获得共鸣。
- MoonshotAI 发布 Kimi-K3 技术报告详述模型架构与训练方法,社区热议自建部署与大公司成本效益,以及对隐私数据托管的态度。
- 迪卡侬德国网站新增基于 SEPA 的 Wero 支付选项,旨在减少对美国支付公司依赖,但部分消费者担忧失去信用卡免息期和争议保护。
- 最新气候模型预测 2026-27 年厄尔尼诺事件可能成为史上最强,峰值远超 2015 年纪录,引发对极端高温应对及全球能耗公平性的讨论。
1. Kimi-K3 在 HuggingFace 上 (Kimi-K3 on HuggingFace) #
https://huggingface.co/moonshotai/Kimi-K3
Kimi K3 是 Moonshot AI 发布的开源多模态代理模型,总参数量 2.8T,激活参数 104B,采用 MoE 架构(896 专家中激活 16 个),结合 Kimi Delta Attention(KDA)和 Attention Residuals(AttnRes)。模型原生支持文本与图像输入,上下文长度达 100 万 token,量化使用 MXFP4 权重和 MXFP8 激活。
在评估中,Kimi K3 在推理、编码、代理任务等多项基准上表现领先,例如 GPQA Diamond 93.5、DeepSWE 67.5、BrowseComp 91.2。页面提供了完整的评估对比表,与 Claude、GPT 等模型相比各有优劣。
使用方式包括通过 Hugging Face Transformers 加载、vLLM、SGLang 等推理框架部署,以及 Docker 模型运行器。模型权重基于 Kimi K3 License 开放。
HN 热度 1301 points | 评论 514 comments | 作者:nateb2022 | 18 hours ago #
https://news.ycombinator.com/item?id=49065752
- Kimi-K3 作为 3T 参数、mxfp4 原生的模型,需要约 1.5TB 显存,托管在 8 块 B200 上勉强够用,但实际可能需要 16 块,成本较高,但可通过 API 定价估算每百万 token 的成本,并判断实验室是否在补贴 API 价格。
- 该模型在 AISI 网络安全基准测试中高于 GLM5.2,但远落后于 SOTA 闭源模型,可能需要微调;也关注 Cursor 是否会用其进行新一轮训练,并与 Kimi2.6/2.7 微调版及 Grok4.5 对比。
- 有人可能尝试从 Kimi-K3 进行完整分布蒸馏到更小的模型(如 DSv4-Kimi),因为 DSv4 服务成本很低。
- 在无 GPU 但配备 1.5~3TB 内存的服务器(如双路/四路 Xeon)上运行该模型很有趣,即使输出速度只有 5-6 tok/s,对长时间任务也可用,且硬件成本远低于 GPU 方案(<3 万美元 vs GPU 硬件)。
- 但电费可能比调用 API 贵 100 倍,只有极端数据隐私需求才可能自建,然而“需要绝对数据隐私、运行 SOTA 模型、买不起 GPU”这三者的交集极窄,甚至可能不存在。
- 美国平均电价远高于 0.075 美元/kWh,各地区差异显著(新英格兰 28.1 美分、中大西洋 25.1 美分、太平洋沿岸 26.1 美分等),但也有部分地区如华盛顿州水电便宜(13.38 美分),或通过太阳能 + 电池、低谷电价等方式降低有效成本。
- 有人实际电费为 0.077 美元/kWh(含税),也有人指出很多人只报低谷电价,而峰值电价更高,实际有效电价会更高。
2. 美国公民因机场搜查期间 GrapheneOS 手机擦除数据被起诉 (US citizen charged after GrapheneOS phone wipes during airport search) #
https://www.techspot.com/news/113236-us-prosecutors-charge-atlanta-man-after-grapheneos-phone.html
美国联邦检察官在亚特兰大起诉一名男子,起因是他使用的隐私保护操作系统 GrapheneOS 在机场搜查时自动擦除了手机数据。当事人 Sam Tunick 从多米尼加共和国返美后,在机场被联邦特工拦截问话,特工怀疑他与反对“Cop City”警察训练设施的运动有关。审讯中,特工要求他解锁手机,他输入密码后手机重启并清空了数据。
检方将此视为故意销毁证据,依据联邦法律提起诉讼;辩方则称搜查违宪,当事人四次要求见律师均被拒绝,且未出示搜查令。案件引发对边境搜查权力与宪法权利的讨论。GrapheneOS 是一款开源系统,旨在提升隐私安全,支持者认为使用该系统本身不应被视为犯罪。法官预计至少要到 10 月底才会对辩方动议作出裁决。
HN 热度 1256 points | 评论 990 comments | 作者:eecc | 1 day ago #
https://news.ycombinator.com/item?id=49063022
- 边境搜查中执法人员权力过大,需要技术手段保护数据,同时避免激怒执法人员
- 技术解决方案可能忽略真实威胁模型,例如汽车 PIN 防盗在绑架场景中反而危险
- 若生物识别等技术普及,可能激励犯罪者绑架或伤害用户以获取访问权限
- 手机防护存在“扳手问题”:只要足够胁迫,任何技术保护都可被绕过,建议携带空白手机过境
- 指纹识别等强制手段虽引发砍手指极端案例,但该事件单一且罕见,并非趋势
- 边境搜查中的胁迫可能是针对抗议者的政治手段,而非真正关心儿童色情内容
- 例行检查与针对性调查应区别对待:例行检查中配合可能更有效,针对性调查中配合无济于事
3. PG 模拟城市 - PostgreSQL 工作原理 (PGSimCity - How PostgreSQL Works) #
https://nikolays.github.io/PGSimCity/
这是一个交互式 PostgreSQL 内部机制模拟器,以 3D 城市地图形式可视化数据库运行状态。每个建筑代表一个核心组件:客户端连接、后端进程、缓冲池(shared_buffers)、WAL、存储、查询实验室、维护(检查点/自动清理)、备用服务器。
用户可以实时调整参数(如 TPS、读写比例、shared_buffers 大小、WAL 级别、同步提交模式等),观察缓存命中率、脏页数量、WAL 速率、复制延迟等指标变化。内置多个预设场景(如检查点风暴、缓存抖动、锁堆积、复制延迟等),帮助理解不同配置下的性能行为。
网站提供 14 步引导教程,从客户端连接到提交完成,逐步讲解每个环节。操作支持鼠标/触摸旋转缩放、键盘快捷键、飞行/步行视角切换。左下角控制台可调节模拟速度、暂停、重置或选择预设场景。
HN 热度 884 points | 评论 86 comments | 作者:jonbaer | 24 hours ago #
https://news.ycombinator.com/item?id=49063754
- 人脑限制与 LLM 辅助构建的细节过载导致体验混乱,动画和隐喻让人困惑,感觉不是为人类设计的。
- 建议减少“Take tour”中的视觉噪音,增加交互性而非自动切换,帮助用户聚焦。
- 3D 效果不错但弹出窗口遮挡了大部分空间,应减少噪声或使用半透明效果。
- 移动端体验不佳。
- 作者通过单个 prompt 和 Claude Opus 5 开始构建,使用了约 3.86B tokens(含缓存),成本因个人计划较低。
- 有用户质疑 token 用量为何如此巨大(对比自己每周 55M tokens)。
- 估算成本约 $19.3k,作者确认通过缓存和个人计划降低。
- 作者分享了初始 prompt:要求构建 3D Postgres 城市模型,类似 city 展示内部组件。
- 有用户喜欢用城市隐喻解释架构,对视觉学习者很有帮助。
- 有人建议在 tour 中添加 TTS(文本转语音)。
- 有人认为需要 tour 说明 UX 有待改进,用户应自然发现功能;但在此场景下 tour 合适,因为不同用户需求不同。
- 期望能输入查询并看到完整流程,当前不知从何开始或结束。
- 按 T 键可能触发 tour,相机旋转帮助在“?”中。
- 赞赏这种直观展示,可复用到其他复杂系统如云计算、Kubernetes 等。
- 有人想用类似 Factorio 的视觉隐喻做部署系统解说器。
4. AI 公司正在粉碎稀有书籍 (AI companies are shredding rare books) #
https://twitter.com/HedgieMarkets/status/2081534588485296565
AI 公司批量收购稀有书籍,使用高速扫描机切割书脊并扫描,随后将原书销毁。服务商 ISBNdb 帮助企业匿名订购多达百万本书籍,并提供保密协议作为特色服务。法官裁定此举属于“合理使用”,因为销毁原书意味着同一时间仅存一份副本。Anthropic 聘请前 Google 图书合作负责人,旨在获取“世上所有书籍”。Hedgie 评论称,这种破坏是不可逆的——历经战火与世纪的稀有书籍被切碎,只为让 AI 学会写更好的营销邮件。Elon Musk 回应表示,已要求 SpaceX AI 团队以不破坏书籍的方式扫描并保存。其他网友将此事比作“亚历山大图书馆 2.0”,认为销毁历史等于允许用 AI 重写历史。
HN 热度 730 points | 评论 462 comments | 作者:anon373839 | 11 hours ago #
https://news.ycombinator.com/item?id=49068738
- 对出版商缺乏同情,他们任由作品绝版直到版权过期,没有理由让受版权保护的书成为稀有书。
- 现代出版质量差,呼吁版权法改革:当出版商无意再版时,其他出版商可免费重印并重新协商作者版税;若只印劣质版本而不印耐久精装版,其他出版商应能印耐久版,无需补偿原出版商。
- 版权应随作者及其配偶去世而终止,希望书籍和音乐版权更宽松,因为唱片公司滥用侵权索赔打压合理使用的视频创作。
- 版权期限应设为 20 年(类似专利),之后作品进入公共领域。
- 固定期限从首次出版算起,避免不知作者去世时间的问题。
- 或者对长期持有版权的公司征收指数增长的税(如迪士尼)。
- 公司已经通过米老鼠赚钱纳税,无论版权状态如何。
- 版权应服务于公共利益,目前做得不够好。
- 20-25 年足够保护艺术创作,其他行业没有这么长的保护,今天出版的作品可能被保护近 200 年。
- 对个人作者而言,20 年期限不公平,会让企业更容易剥削作者;版权与专利不同,专利有直接实用价值,且衍生作品的价值有限(反驳观点认为开源证明衍生作品有价值)。
5. Bun 用 Rust 重写进展如何? (How is the Bun rewrite in Rust going?) #
https://lockwood.dev/ai/2026/07/27/how-is-the-bun-rewrite-in-rust-going.html
开发者 Tom Lockwood 在 2026 年 7 月 27 日发表文章,质疑 Bun 团队声称的“用 Rust 重写”的真实性与成本。文中指出,Bun 创始人 Jarred Summner 称在 11 天内花费 16.5 万美元使用 Anthropic API 完成重写并合并到主分支。然而自合并至今 6 周,仍无发布标签,上次发布已是 11 周前。作者克隆仓库后,发现机器人 robobun 创建的 PR 数量从 1277 激增至 2475,合并需持续运行 CI/CD 超过 86 天。此外,Athropic 员工深度参与,实际花费可能远高于公布的 16.5 万美元,接近 80 万美元。作者认为支持公司估值的声明需要谨慎看待,并指出 Anthropic 的 C 编译器和 Cursor 的 FastRender 浏览器已数月无提交。
HN 热度 446 points | 评论 339 comments | 作者:tomlockwood | 13 hours ago #
https://news.ycombinator.com/item?id=49067854
- 用 Rust 重写 Bun 进展顺利,已在 Claude Code 中发布一个多月,但 v1.4 因需提升 Node.js 测试兼容性而延迟,预计下周二发布
- 花时间保证软件质量没问题,一个月不发布不算大事,批评文章毫无根据
- 最初“10 天和 16.5 万美元”的估计过于乐观
- 作者并非要求发布,只是与同行聊天时查看进展,对新技术的怀疑态度是合理的
- CI 一直很贵,Bun 跨平台构建并多机器分片测试,最近改进了交叉编译,每日用 Claude 优化慢测试
- 关于 CI 成本“80 万美元 +”的估算是否准确?Jarred 没有直接回应
- Bun 使用 BuildKite 而非自托管 CI,因为 BuildKite 能按需启动临时实例,比 GitHub Actions 便宜,未来可能会切换到自定义方案
- 每月 CI 花费是多少?随着 robobun 的使用可能大幅增加
6. 法国消防员首次面临“火积云” (French firefighters face ‘pyrocumulonimbus’ for first time) #
法国西南部一场巨大的野火首次在该国产生了极其危险的“火积云”现象(又称积雨云 flammagenitus),此前这种现象仅在野火多发的澳大利亚和北美出现过。法国国家消防队联盟发言人埃里克·布罗卡迪中校解释,地面温度极端升高,热空气急速上升并与上层冷空气相遇,形成了缺乏水分的“火云”,它生成自己的天气系统,使路径上的任何物体自燃,云内还会产生闪电,进一步助长火势。
这种“对流性”火灾会不断改变风向,向各个方向蔓延,形成多个火线,完全无法预测,消防员无法直接对抗,如同大卫对战歌利亚,只能寻找薄弱点攻击。控制火势的希望要么来自持续三天的暴雨,要么将火引向大海等自然消亡的地方。
消防员已进入“作战不可能”状态,必须战略撤退,承认自然力量的不可控。现场形势是防御性的,类似战争中的持续轰炸,首要任务是确保人员安全。尽管感到无力,但消防员不会绝望,每一项行动都在为最终结果注入活力,但全国消防力量不能全部集中于此,因为风险无处不在。
HN 热度 444 points | 评论 355 comments | 作者:saaaaaam | 1 day ago #
https://news.ycombinator.com/item?id=49060495
- 火积云在法国并非首次出现,但这次强度和频率前所未有。
- 气候灾难的强度和频率在明显增加,而人们只争论云的形状,忽视了根本问题。
- 讨论云的形状体现的是对准确性的追求,并非功能失调。
- 大火能产生自己的雷暴是火灾强度的直接体现。
- 大部分火灾是人为引起的,而非自然,气候变化不一定是主因。
- 气候变暖导致森林干燥,更容易点燃,人为点火只是导火索。
- 森林干燥和闪电增多是火灾频发的主因。
- 纵火不否定气候变化影响,但需要全面研究环境因素和预防措施。
- 质疑气候变化叙事需要具体模型和数据支撑,而非简单归因。
7. Decker:一个继承 HyperCard 和经典 macOS 传统的平台 (Decker, a platform that builds on the legacy of Hypercard and classic macOS) #
https://beyondloom.com/decker/
Decker 是一个多媒体平台,用于创建和分享交互式文档,支持声音、图像、超文本和脚本行为。它继承了 HyperCard 的简洁性和经典 MacOS 的视觉美学,并加入了深度撤销历史、滚轮/触屏支持、现代键盘导航等改进。用户可用它制作电子杂志、整理笔记、演示、冒险游戏或绘制 1-bit 像素画。成品可保存为独立的 HTML 文件,在浏览器中运行,并可在任意平台托管或嵌入。
Decker 内置一门名为 Lil 的脚本语言,融合了 Lua 的实用性与 Q 的函数式特性,支持隐式标量-向量运算和内嵌 SQL 查询。它还提供少量内置交互小部件,并允许用户定义新部件,通过剪贴板复制粘贴共享。Decker 命令行友好,附带的 Lilt 解释器可独立运行,并能“无头”读写、操作和执行 Decker 文档。所有文档采用行文本格式,与 Git/SVN 等版本控制工具良好协作。
Decker 不含广告、遥测、游戏化或隐私侵犯。它遵循 MIT 开源许可,源代码和问题追踪在 GitHub 上,二进制版本在 Itch.io 发布,社区论坛和每两年一次的“游戏创作节”也在此进行。页面还提供了大量示例、库(如绘图、统计、PDF 导出、动画等)以及详细文档(手册、Lil 语言指南、格式说明等)。
HN 热度 374 points | 评论 99 comments | 作者:tosh | 1 day ago #
https://news.ycombinator.com/item?id=49060856
- HyperCard 提供了非凡的体验,即使对儿童也直观易用,能制作真实应用,且看起来不廉价。
- “看起来像软件”的能力在图形界面时代更难实现,HyperCard 和 Visual Basic 是很好的中间地带。
- AI 作为 UI 提供了类似的专业感,可以快速构建功能并显得专业。
- 限制激发创造力,《辐射》类比、demoscene 等例子说明约束能培养创意。
- 真正的美来自在约束中表达,现代艺术中出现了回归色彩和多样性的趋势。
- 直到 2010 年浏览器才匹配 HyperCard 的交互性,但网络才是杀手应用,事件驱动开发模式后来才普及。
- 《神秘岛》等游戏是用 HyperCard 制作的,其影响力巨大。
- 只需让人们尝试使用 HyperCard 就能体会其魅力,现代 LLM 可编写 Hypertalk。
- 2D 等距游戏的美感并非仅因低分辨率,而在于像素比例和拼接的精确性。
- 类似 HyperCard 的工具(如 Access、FileMaker)需求巨大,但被抛弃,现在小企业数据库缺乏默认选择,Notion 可能是一种替代。
8. Kimi-K3 技术报告 [pdf] (Kimi-K3 Technical Report [pdf]) #
https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf
MoonshotAI(月之暗面)在 GitHub 上公开了 Kimi-K3 项目的技术报告(k3_tech_report.pdf),这是一个名为 Kimi-K3 的模型,目前该仓库获得 1.1k 星标和 83 个 fork。报告中可能包含模型架构、训练方法、性能评估等关键技术细节,旨在推动 AI 开源社区的发展。
HN 热度 361 points | 评论 162 comments | 作者:vinhnx | 9 hours ago #
https://news.ycombinator.com/item?id=49070985
- 对于每年推理费用超过百万美元的大公司,购买一台 GB300 机架(约 600 万美元)来运行 Kimi-K3 模型是划算的,年摊销加电费约 150 万美元,每百万输出 token 成本低于 60 美分,远低于 API 调用费用。
- 有用户正在用 10.7 万美元的服务器运行 Kimi 2.8,每秒处理约 5 万 token,相比去年使用 Claude/Gemini 节省了成本。
- 运行该机架需要雇佣 2-3 名运维人员,每年额外增加 40-70 万美元成本,且将问题从厂商转移到了自己身上。
- 对于这类大公司,通常已有专职运维团队,添加一个机架不需要额外招聘,工作负载极少。
- 许多公司不信任托管的 LLM 处理隐私数据,宁愿自建,但也有人指出 AWS 等云服务可能更可靠。
- 即使自建模型,也无法完全防止模型在应用中被嵌入恶意二进制内容,除非从头创建 LLM。
9. 迪卡侬德国在 decathlon.de 网站新增 Wero 支付选项 (Decathlon Germany adds Wero payment option to decathlon.de website) #
该网页是一个网站 Cookie 隐私政策与偏好管理界面,主要说明网站使用 Cookie 以增强浏览体验、提供个性化广告及分析流量。页面将 Cookie 分为三类:
- 必要 Cookie:用于实现网站基本功能,如安全登录、会话管理和反伪造令牌,例如
WV_SESSION、ASP.NET_SessionId、__cf_bm等。 - 功能 Cookie:支持社交媒体内容分享、收集用户反馈及其他第三方功能,如 LinkedIn 的
lidc、YouTube 播放器设置相关的yt-remote-*系列 Cookie。 - 分析 Cookie:帮助网站了解访客如何互动,统计访问量、跳出率等,包括 Google Analytics 的
_ga、_gid、HubSpot 的hubspotutk等。
每个 Cookie 都列出了名称、持续时间(如会话级或 1 年)及简短作用描述。用户可通过“Accept All”接受全部 Cookie,或点击“Customize Consent Preferences”进行个性化设置,随时管理或撤回同意。
HN 热度 296 points | 评论 207 comments | 作者:doener | 7 hours ago #
https://news.ycombinator.com/item?id=49072310
- Wero 基于 SEPA 系统,提供银行转账的安全性和便利性,无需账户且不受美国公司控制。
- Wero 基于荷兰的 iDEAL 系统,支持 P2P 支付和 QR 码,计划迁移至 Wero 并支持非接触支付以减少对 Google/Apple Pay 的依赖。
- 荷兰法律规定在线商店不能只接受全额预付,因此多数商店同时支持信用卡,信用卡满足 50% 后付款规则并提供消费者保护。
- 信用卡提供免息 30 天贷款和争议解决,对消费者有利;若商家拒绝信用卡改用 Wero,可能降低购买意愿。
- 欧洲使用借记卡为主,消费者债务多为房贷,与美国的消费信贷情况不同。
- Wero 目前只是 iDEAL 2.0 更换 Logo,技术未变,实际迁移从 10 月开始。
- Wero 是 Google/Apple 生态独占,而 iDEAL 支持银行网站支付。
10. 史上最强厄尔尼诺 (The Strongest El Niño Ever) #
https://www.theclimatebrink.com/p/the-strongest-el-nino-ever
根据 14 个季节性预测模型的最新 7 月数据,2026-27 年厄尔尼诺事件极有可能成为有可靠记录以来最强的一次。模型集合中位数预测峰值(Niño 3.4 区海表温度异常)达到 3.6°C,远超此前 2015-16 年 2.75°C 的纪录,差距约 0.8°C。约 91% 的集合成员超过 2015-16 年纪录,即使使用去除长期趋势的相对指数(RONI),也有 77% 的成员创下新高。事件发展速度超过 1997-98 年,且从年初拉尼娜条件迅速演变。各模型高度一致,所有模型预测均达到“超级厄尔尼诺”强度。当前 7 月中旬实测值已接近 2°C 异常,若维持则将达到超级厄尔尼诺阈值。
HN 热度 280 points | 评论 336 comments | 作者:ndsipa_pomu | 1 day ago #
https://news.ycombinator.com/item?id=49060978
- 建议欧洲人安装太阳能和空调系统应对 2027 年创纪录高温。
- 批评美国自身高能耗却嘲笑欧洲缺乏空调的双重标准。
- 欧洲部分国家对公共建筑空调设置最低温度限制,但近年安装量逐渐增加。
- 欧洲过去高温天数少,如今热浪持续数周,热死亡风险高,尤其老年人。
- 欧洲热死亡统计可能使用超额死亡率,与美国直接归因统计方法不同。
- 欧洲社会保守,对改变空调使用习惯持抵制态度。
- 西班牙南部广泛使用空调且依靠可再生能源,瑞士等国对固定空调有安装限制。
- 移动空调效果有限,难以应对持续高温。
- 欧洲是全球气温上升最快地区,但适应措施不足。
- 老年人口在热浪中死亡率极高,养老院等公共设施缺乏空调是主要原因。
Hacker News 精彩评论及翻译 #
Our position on open-weights models #
https://news.ycombinator.com/item?id=49076261
Anthropic has never advocated for a ban on open-weights models.
All sufficiently capable models, open and closed, should go through mandatory safety testing.
Yeah, this is anthropic advocating for a ban on open weight models.
Who runs this test? What happens if this test is too costly or the administrator refuses to allow certain people to participate.
This is exactly how the US has banned goods in the past, by requiring a stamp and then refusing to issue it.
cogman10
Anthropic 从未主张禁止开源权重模型。
所有足够强大的模型,无论是开源还是闭源,都应进行强制性安全测试。
对,这就是 Anthropic 在主张禁止开源权重模型。
谁来执行这些测试?如果测试成本过高,或者管理者拒绝让某些人参与,又该怎么办?
这正是美国过去禁止商品的方式:先要求某种认证,然后拒绝颁发。
Kimi-K3 on HuggingFace #
https://news.ycombinator.com/item?id=49065868
This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it’s going to be mxfp4 native, it’ll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you’ll need 16x for context / throughput optimisation). Won’t be cheap to host, but at least we should get some range of $/MTok for a 3T model. Then we’ll be able to guesstimate if “labs are subsidising tokens on API pricing”.
Also interesting to see what effort it will take to fine-tune this beast. The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models. Some fine-tuning might be needed here. Also, interesting to see if Cursor does another training round on it, to directly compare it w/ kimi2.6/2.7 fine-tunes (composer series) and grok4.5.
Also also, interesting to see if someone takes on distilling (proper distillation, w/ training the entire distribution) from this into smaller models. (dsv4-kimi should be really good, since dsv4 is very cheap to serve)
NitpickLawyer
这将会很有趣,原因有几方面。首先,第三方提供商的中位定价将告诉我们运行一个3T模型的实际成本。由于它原生采用mxfp4格式,大约需要1.5TB显存来部署——这刚好处于8块B200的极限(但实际为了上下文/吞吐量优化,你需要16块)。托管成本不会低,但至少我们能了解到3T模型每百万token的美元价格区间,从而推测“实验室是否在补贴API定价中的token费用”。
另一个有趣的点是,微调这个庞然大物需要多少努力。最新的AISI网络安全基准测试显示,它高于glm5.2,但仍远落后于最先进的闭源模型。可能需要一些微调。此外,值得关注的是Cursor是否会对其进行新一轮训练,以便直接与微调后的kimi2.6/2.7(composer系列)以及Grok4.5进行对比。
同样有意思的是,是否会有人从这个模型中蒸馏出更小的版本(真正的蒸馏,训练整个分布)。dsv4-kimi应该会很不错,因为dsv4的部署成本非常低。
How is the Bun rewrite in Rust going? #
https://news.ycombinator.com/item?id=49069787
Bun’s Rust rewrite shipped in Claude Code over a month ago and barely anyone noticed. Claude Code is widely used. The Rust rewrite is going well overall.
In the Bun v1.4 video, I promised a certain number of newly passing Node.js tests were added to force us to improve compatibility, and that number is not true yet. The release is delayed until it is true. The PRs to make it true are up but not merged yet. Most likely next Tuesday we’ll do the release of 1.4.
Jarred
Bun的Rust重写版在超过一个月前就随Claude Code发布了,但几乎没人注意到。Claude Code被广泛使用,Rust重写整体进展顺利。
在Bun v1.4视频中,我承诺会新增一定数量的通过Node.js测试,以强制我们改进兼容性,但目前这个数字尚未达成。发布将推迟到该目标实现。相关PR已经提交但尚未合并。很可能下周二我们会发布1.4版本。
US citizen charged after GrapheneOS phone wipes du… #
https://news.ycombinator.com/item?id=49064242
I’ve seen a lot of people on the internet over the years say things like “the government can’t make x illegal, it’s just y.” For example, the government can’t make wiping your phone at the border illegal, it’s just punching four numbers into your phone, just like a pin, only a different four numbers, which could just have well been your pin.
U.S. law though is highly non-autistic and what you were trying to do is just as important as what you superficially did. Hell there could have been a third set of four numbers that were the nuclear launch codes. It’s not the fact that it was four numbers, it’s what you were trying to make happen when you typed them. Now of course whether they can prove what your intent was when you typed them is another matter, but generally a duress pin should be for when robbers are breaking into your house, and the government will be on your side, and not when the government will be against you.
cameldrv
多年来我在网上看到很多人说类似“政府不能把X定为非法,它只不过是Y”的话。比如,政府不能把在边境擦手机定为非法,它只不过是往手机里输入四个数字,就像密码一样,只是不同的四个数字,可能正好是你的密码。
但美国法律非常不刻板,你试图做的事情和你表面上做的事情同样重要。见鬼,甚至可能有第三组四个数字是核发射密码。关键不在于它是四个数字,而在于你输入它们时试图达成什么目的。当然,他们能否证明你输入时的意图是另一回事,但一般来说,胁迫密码应该用于强盗闯入你家时,那时政府会站在你这边,而不是当政府与你为敌时。
US citizen charged after GrapheneOS phone wipes du… #
https://news.ycombinator.com/item?id=49066003
I co-wrote a border search guide for EFF some years ago. I was very interested in finding clever technical approaches but I later ended up feeling that I hadn’t given enough thought to the overall threat model questions (even though the guide did address them, perhaps even somewhat usefully).
The big picture problem is that the agents performing the searches have an enormous amount of power in terms of potentially seizing devices and potentially denying entry for non-citizens. I think they should not have this power, but the agents and courts probably don’t care that I think that.
The end result (not inherently different from what we wrote in the guide) is that you may have to think both about protecting your data by technical means, and about not angering the agents more than you plan to. I was fascinated by techniques for being unable to comply (which is straightforward to achieve if you want!) but probably didn’t think enough about how much this might antagonize border agents in many cases.
I definitely don’t know a comprehensive big-picture solution.
schoen
几年前,我与他人合作为电子前哨基金会(EFF)撰写了一份边境搜查指南。当时我对寻找巧妙的技术方法非常感兴趣,但后来我意识到自己并没有充分思考整体的威胁模型问题(尽管指南确实涉及了这些问题,或许甚至有些用处)。
宏观问题在于,执行搜查的探员拥有巨大的权力,可以扣押设备,并可能拒绝非公民入境。我认为他们不应拥有这种权力,但探员和法院可能并不在意我的看法。
最终结果(与我们在指南中写的本质上并无不同)是,你可能需要同时考虑通过技术手段保护数据,以及避免比计划中更激怒探员。我曾痴迷于如何做到无法配合的技术(如果你想做到,这很简单!),但可能没有充分思考这在很多情况下会如何激怒边境探员。
我确实不知道一个全面的宏观解决方案。
French firefighters face ‘pyrocumulonimbus’ for fi… #
https://news.ycombinator.com/item?id=49065091
The main reason the forest is burning is global warming. There were 3 punishing heatwaves before the fire, it was just a matter of time.
A part of the Fontainebleau forest, much more north, no pine trees, went in smoke earlier. Forests are burning near Madrid.
It may make some people feel better to think that it is an isolated, local problem, but unfortunately global warming is the main culprit and we at least need to acknowledge it.
Article explaining when pine tree forest resist well or not to fire. https://www.sudouest.fr/faits-divers/incendies/incendies-en-gironde-et-dans-les-landes-le-probleme-n-est-pas-le-pin-mais-la-secheresse-30061837.php
reco98gg
森林燃烧的主要原因是全球变暖。火灾发生前经历了三次酷热的热浪,起火只是时间问题。
更靠北的枫丹白露森林一部分——那里没有松树——更早之前就化为了灰烬。马德里附近的森林也在燃烧。
认为这只是个孤立、局部的问题或许能让一些人感觉好受些,但不幸的是全球变暖才是罪魁祸首,我们至少需要承认这一点。
文章解释了松树林在什么情况下能有效抵御火灾,什么情况下不能。https://www.sudouest.fr/faits-divers/incendies/incendies-en-gironde-et-dans-les-landes-le-probleme-n-est-pas-le-pin-mais-la-secheresse-30061837.php
French firefighters face ‘pyrocumulonimbus’ for fi… #
https://news.ycombinator.com/item?id=49063527
For context on why this region is burning so easily: the Landes and Médoc are huge, artificial pine forests created in the 19th century under Napoleon III to drain and reclaim what was considered an inhospitable wetland, where shepherds famously walked around on stilts. Pine resin and needle litter makes it exceptionally flammable, and being an uninterrupted monoculture on flat land it lacks natural barriers to stop the fire from advancing.
Dibby053
为了解释为何这片区域如此容易燃烧:朗德和梅多克地区是19世纪拿破仑三世时期人为种植的广阔松林,旨在排干并开垦当时被认为不宜居住的湿地,而当地的牧羊人过去常踩着高跷行走。松脂和针叶落叶使其极其易燃,加上平坦土地上连绵不断的单一作物林缺乏阻止火势蔓延的天然屏障。
AI companies spend record sums on Washington lobby… #
https://news.ycombinator.com/item?id=49070629
OpenAI nearly doubled its federal lobbying expenditure to a record $2.22mn in the first half of 2026, compared to last year, while Anthropic nearly tripled its spending to $3.53mn, according to federal disclosures.
Never ceases to amaze me how cheap lobbying is. That’s pocket change for these companies.
simonw
游说活动如此便宜,这总是让我感到惊讶。对这些公司来说,这只是零花钱。
Should you wash your solar panels? #
https://news.ycombinator.com/item?id=49069645
why would you clean them when you’re about to sell the house?
So they look good and the potential buyer sees something that looks to be in good condition and working well.
wodenokoto
为什么你都要卖房子了还要清理它们?
这样它们看起来状态好,潜在买家看到的是保养得当、运转正常的东西。
French firefighters face ‘pyrocumulonimbus’ for fi… #
https://news.ycombinator.com/item?id=49060584
The situation in Bordeaux right now feels close to apocalyptic. I left yesterday. Two hundred thousand people are evacuated, hundreds of homes destroyed, and the fire about 10 miles from the edge of the city.
verzali
波尔多当前的局势近乎末日。我昨天离开了。二十万人被疏散,数百座房屋被毁,大火距城市边缘约10英里。
US citizen charged after GrapheneOS phone wipes du… #
https://news.ycombinator.com/item?id=49064789
Federal prosecutor success rate is over > 90%.
This is a misunderstood statistic.
Federal prosecutors won’t even pursue cases unless they think there’s a high chance of success. They don’t operate like two private parties suing each other to force the court to decide something. If the evidence is there or the charges aren’t fully formed, they don’t waste resources on it.
This leads to a contradictory set of complaints that the legal system lets too many people go or doesn’t have enough teeth.
Aurornis
联邦检察官的成功率超过90%。
这是一个被误解的统计数字。
联邦检察官只有在认为胜算很高时才会提起诉讼。他们不像私人双方那样通过诉讼迫使法院做出裁决。如果证据不足或指控尚未完全成立,他们不会浪费资源。
这导致了一组矛盾的抱怨:要么是法律系统放走了太多人,要么是它不够强硬。
Kill The Cookie Banner #
https://news.ycombinator.com/item?id=49060371
“Accept all” always makes it go away immediately.
Some variant on “reject” takes more effort like 70% of the time. Which is on purpose, of course. The ones that aren’t maliciously-complying have a “necessary only” button that insta-closes it, but tons pretend that you might want to allow some spying but not all of it and make you go through another screen if you don’t just “accept all”.
To me, having a browser setting for cookies is the only sane way to handle this, it’s surprising that this was not considered from the beginning.
Then it’d be possible to default it to “nope” (Firefox, and perhaps Safari, might do this) or to allow a “never, anywhere” setting the first time the question is asked, and malware and spyware vendors know that’d mean a much larger proportion of denials.
readread
“接受全部” 总是能立刻让它消失。
而类似于"拒绝"的选项,大概有70%的情况下需要花费更多功夫。这当然是故意的。那些不是恶意顺从的网站会有一个"仅必要"按钮,可以立即关闭弹窗,但大量网站会假装你可能想允许某些追踪而非全部,如果你不点击"接受全部",就会让你进入另一个界面。
对我来说,浏览器设置一个针对Cookie的选项才是处理这个问题的唯一合理方式,令人惊讶的是这一点从一开始就没有被考虑。
这样一来,就可以默认设为"拒绝"(火狐浏览器,或许还有Safari,可能已经在这样做),或者在第一次询问时允许设置"永远、任何地方都拒绝"的选项,而恶意软件和间谍软件供应商知道,这意味着拒绝的比例会大得多。
EU Fines Google $1.02B for Favoring Its Own Servic… #
https://news.ycombinator.com/item?id=49066209
The EU told them they can not by default prefer their own maps product when linking from their search product, they have to allow the user to choose.
They could have simply added a selector when a user first clicks on the maps preview in the search result, and then remembered it on device or across that user’s account.
But of course, then the user could choose a competitor’s product and Google would have to honor it. That would be horrible, so instead they just made the user experience worse for everyone and elegantly made people blame the EU.
phiresky
欧盟告诉他们,从搜索产品链接时不能默认优先使用自己的地图产品,必须让用户自行选择。他们本可以在用户首次点击搜索结果中的地图预览时添加一个选择器,然后在设备上或跨用户账号记住这一选择。但当然,这样一来用户可能会选择竞争对手的产品,而谷歌必须遵守。那太可怕了,所以他们反而让所有人的用户体验变得更糟,并巧妙地让人们把矛头指向欧盟。
It’s not empowering to hand off the details #
https://news.ycombinator.com/item?id=49062508
I’ve been vibecoding a ton for the past 9 months, built a bunch of cool little apps for myself with AI, ran experiments, built an entire SDLC on skills, did the agent orchestration harness thing, etc. In the past few weeks I’ve hit a wall where I’m just tired of it. Each model becomes more independent but also harder to direct in detail. They produce massive, tedious, sloppy text outputs with very little input. They’re bad at socializing knowledge and communicating design forks.
The places where I’ve seen unequivocal wins with AI are repetitive tech debt tasks that apply the same transformation across a large amount of code or refactor under a pre-existing test suite with good coverage. It’s great for initial research, brainstorming, and can be good (despite the sycophancy) as a rubber duck conversation partner. I use AI constantly, for work and in my personal time, but we’ve hit a ceiling where I no longer find it helpful for the models to absorb more of the intellectual labor. They get things wrong more aggressively, and more elaborately. They’re inadequately curious. I cannot keep up with the endless bad technical writing, and it makes it harder to spot factual errors and bad reasoning.
Here’s what I want: I want AI as an assistant that helps me make decisions, and ensures that I’m in the driver’s seat. AI as an over-confident prodigy on speed is what we’re getting lately, and it’s losing me.
RGS1811
过去九个月我沉迷于"振动编码",用AI给自己造了一堆酷炫的小应用,做了各种实验,构建了一整套基于技能的系统开发生命周期,还搞了智能体编排框架之类的东西。但最近几周我撞上了墙——纯粹是厌倦了。每个模型都越来越独立,却也越来越难精准指挥。它们用极少的输入就能生成海量冗长、潦草又粗糙的文本输出,不善于传递知识,也不擅长沟通设计分支。
在我看来,AI明确胜出的领域是重复性的技术债务任务——比如在大段代码中执行统一转换,或在已有高覆盖率测试套件下进行重构。它对前期调研和头脑风暴很有帮助,即便有些谄媚讨好,作为"橡皮鸭"式的对话伙伴也不错。我无时无刻不在用AI,无论是工作还是个人时间,但我们显然撞上了天花板——模型吸收更多脑力劳动这件事,对我而言已不再有益。它们犯错的姿态更强势、更煞有介事,却缺乏足够的好奇心。我无法忍受无休止的劣质技术写作,这让我更难发现事实错误和逻辑谬误。
我真正想要的是:AI成为辅助我做决策的助手,确保我始终掌控主导权。但最近我们得到的却是"嗑了兴奋剂的自大天才"——这让我逐渐失去兴趣。
US citizen charged after GrapheneOS phone wipes du… #
https://news.ycombinator.com/item?id=49064519
“U.S. law though is highly non-autistic” hilarious but also another point to emphasize is how truly depressing American courts often are. Take the right to a jury. It sounds noble in theory. But when they say judged by your peers they don’t mean your actual peers.
It’s people who couldn’t get out of jury duty. Prosecutors have high success rates. Federal prosecutor success rate is over > 90%. Studies of jury psychology show how much peer pressure and other factors extrinsic to the law come into play.
Remember what happened to Aaron Swartz. Law is the mask of power. By all means defend and assert your rights, but understand the costs. I find people are under such illusions about how cruel the American justice system is that this leads them to make foolish decisions. Do not underestimate the adversarial nature of the justice system, nor the accompanying incentives agents of the state who are on the other side of you have to lie.
godwinson__4-8
“美国法律虽然非常不适于自闭症患者” 这话很搞笑,但另一点需要强调的是,美国法庭实际上往往令人沮丧。以陪审团权利为例,理论上听起来很崇高。但当他们说"由你的同龄人审判"时,他们指的并不是你真正的同龄人。
那些是没能逃脱陪审义务的人。检察官胜诉率很高,联邦检察官胜诉率超过90%。关于陪审团心理的研究显示,同侪压力以及其他法律之外的因素影响有多大。
还记得亚伦·斯沃茨的遭遇吧。法律是权力的面具。务必捍卫和主张你的权利,但要理解代价。我发现人们对美国司法系统的残酷程度存在如此大的错觉,以至于他们做出愚蠢的决定。不要低估司法系统的对抗性,也不要低估与你对立的国家代理人随之而来的说谎动机。
Decathlon Germany adds Wero payment option to deca… #
https://news.ycombinator.com/item?id=49073093
Wero is such a valuable asset for Europe and couldn’t have come at a better time.
It’s built on the excellent SEPA system, which standardized bank transfers and “Lastschrift” in Europe. SEPA itself got a huge upgrade last year with mandatory instant transfers that cannot cost more than standard transfers between European checking accounts.
Both of those allowed Wero to happen as a final abstract top layer, meaning: the UX of sending money to an email address with the security and convenience of bank transfers. And everything is built collaboratively on top of European banking infrastructure.
No account necessary and no private american conglomerates can delete that account.
And before some says: “Well, now the government controls everything!”: They already did, since… Ever. Now its one less entity with that access and one that can be (hopefully) controlled and changed by a majority.
anticrymactic
Wero 对欧洲来说是极为宝贵的资产,而且来得正是时候。
它建立在卓越的SEPA系统之上,该系统标准化了欧洲的银行转账和"直接借记"业务。去年,SEPA 本身也获得了重大升级,强制要求即时转账,且费用不得超过欧洲活期账户之间的标准转账。
这两点使得Wero能够作为最终的抽象顶层得以实现,这意味着:通过电子邮件地址转账的体验,兼具银行转账的安全性和便利性。而这一切都是在欧洲银行基础设施之上协作构建的。
无需开设账户,也没有美国私人企业集团能删除这个账户。
如果有人要说:“好吧,现在政府控制了一切!"——他们早就控制了一切,从……一直以来都是如此。现在只不过少了一个拥有这种访问权限的实体,而这个实体(希望)能够被多数人控制和改变。
MAI-Cyber-1-Flash inside MDASH #
https://news.ycombinator.com/item?id=49073183
Cybersecurity is not just a data-rich domain; it is a live reinforcement learning loop.
twitch
eat
网络安全不仅仅是一个数据丰富的领域;它是一个实时的强化学习循环。
US citizen charged after GrapheneOS phone wipes du… #
https://news.ycombinator.com/item?id=49063972
Ultimately, when you choose to enter a duress PIN that will wipe your device, you have to recognize that choice may have legal consequences. I don’t like the amount of power our government has at the national border when it comes to detaining and pressuring citizens, but our Constitution explicitly grants it at least some of the power it now exercises in that context.
If your threat model includes US state actors at the national border, then your security practices need to account for the confiscation of your device at that border without requiring you to willfully wipe the phone and (in the eyes of police and prosecutors) destroy evidence.
That means:
-
Don’t travel with anything you can’t afford to lose on device. This means setting up travel-specific password managers and hardware keys for a subset of your accounts that you absolutely need to access while abroad, and being prepared to reset those passwords and disable those hardware keys very quickly once home.
-
Review past legal cases against travelers and identify what behaviors the government considers worthy of prosecution or harassment. Your secure setup must function without needing you to engage in those behaviors , even if it is less convenient as a result. This isn’t perfect, as the government may decide some new behavior is prosecutable.
-
Consult with a lawyer and review your security procedures from a legal standpoint. All of the above is technical and practical advice, not legal counsel and no substitute for it.
We Americans are fortunate to carry powerful passports and enjoy relatively easy international travel but, for better or worse, that velvet glove covers an iron fist we would be foolish to forget or ignore.
sfRattan
最终,当你选择输入一个会清除设备的胁迫密码时,你必须认识到这一选择可能带来法律后果。我不喜欢我们的政府在边境上对公民进行拘留和施压时所拥有的权力,但美国宪法明确赋予了它至少部分当前在此情境下行使的权力。
如果你的威胁模型包括美国边境上的国家行为者,那么你的安全实践需要考虑在边境上设备被没收的情况,而无需你故意清除手机,并且在警方和检察官眼中,这被视为销毁证据。
这意味着:
-
不要携带任何你无法承受丢失的设备内容出行。这意味着为你出国期间绝对需要访问的部分账户设置旅行专用的密码管理器和硬件密钥,并准备好回国后迅速重置这些密码并禁用那些硬件密钥。
-
回顾过去针对旅行者的法律案例,识别政府认为值得起诉或骚扰的行为。你的安全设置必须能够在无需你参与这些行为的情况下正常运行,即使这样会带来不便。这并非完美,因为政府可能会将某些新行为认定为可起诉的。
-
咨询律师,并从法律角度审查你的安全流程。以上所有内容均为技术和实践建议,而非法律咨询,也不能替代法律咨询。
我们美国人幸运地持有强大的护照,享受相对便利的国际旅行,但无论好坏,这层天鹅绒手套之下包裹着一只铁拳,我们若忘记或忽视它便是愚蠢的。
Htmx 4.0, the first JavaScript library to release … #
https://news.ycombinator.com/item?id=49059377
I’ve been using htmx for three years now, and it’s completely unlocked new ways of building software for the web, especially if you use a server-side templating language.
But more importantly (as this “Game Boy” option reveals), the head honcho is really responsive with the items in the store. I had purchased a coffee mug a few years back and complained on Twitter that they need to sell larger mugs instead of only offering little tiny baby mugs that only hold ~8 ounces of liquid. The next day, he added giant 48-ounce monsters to the store and I’ve been drinking out of that mug every day since.
blister
我用htmx已经三年了,它完全解锁了构建网络软件的新方式,尤其是当你使用服务器端模板语言时。
但更重要的是(正如这个"Game Boy"选项所揭示的),大老板对商店里商品的反馈非常迅速。几年前我买了一个咖啡杯,然后在推特上抱怨说他们应该卖更大的杯子,而不是只提供那些只能装大约8盎司液体的小不点杯子。第二天,他就在店里添加了巨大的48盎司巨无霸杯,从那以后我每天都用那个杯子喝水。
Show HN: Physically accurate black hole you can pu… #
https://news.ycombinator.com/item?id=49065148
If anyone’s interested in the accuracy, this is very ‘visualisation grade’ software. Its a bit of a pet peeve of mine that people present these sims as being very physically accurate, when they contain major inaccuracies, some of which are very obvious and/or deliberate. This one has some serious physical limitations
I wouldn’t mind at all if it didn’t say that this was a physically accurate black hole specifically, but this now falls under misleading science communication in a way that often gets hand waved away as if it doesn’t matter, so we’ve got to clear some things up!
-
The accretion disk shouldn’t be red, people just expect it because it looks cool. Black hole accretion disks are near universally hot enough to be blue. Interstellar did this too, and tried to handwave it away very unconvincingly
-
This is a non/low (?) spin black hole, which isn’t super duper realistic
-
It ignores the position of the camera (which affects the lorentz shifting)
-
The doppler shifting isn’t terribly accurate
-
It doesn’t model the accretion disk temperature distribution or colour with any kind of accuracy. Usually you model accretion disks as a blackbody radiator, shift it by the doppler, and to display this convolve this against the human eye response (LMS), go to XYZ, then RGB, do a physical tonemapping step, before an sRGB conversion. This instead does none of that - no step of this is done with any physical accuracy. Its not even illustratively correct as we’ll get into
-
The actual radiative transfer is very simplified compared to what you’d use for realsies, and isn’t based on any real numbers, with very simplified equations. The opacity and emissivity of the disk is arbitrary, as is the size, and it does not correctly incorporate brightness or extinction, eg here https://github.com/aplavin/blackhole.plav.in/blob/2f004bfeca670d4dfdaf24dc5c34af862f5f867b/src/raymarch.frag#L379 is super simplified
-
The wrong equation is used for the doppler calculation. They use the I^3 variant, whereas the data you get out of a disk sample is radiant flux which is actually F_obs = F_emit / (z+1)^4. This is a very common mistake in image processing, which means that the doppler shift and observer brightness isn’t correct. Surface brightness over here https://github.com/aplavin/blackhole.plav.in/blob/2f004bfeca670d4dfdaf24dc5c34af862f5f867b/src/raymarch.frag#L47 is not a spectral radiance but instead a radiant flux
Stuff like the brightness -> colouring conversion is particularly inaccurate. Eg if you check out the source:
It maps the pseudo brightness completely arbitrarily to colour. The resulting colour/brightness here then doesn’t correspond to anything remotely physical. It also performs a linear mapping of a linear quantity (brightness) to sRGB (which is a nonlinear process!!), which means that it doesn’t even retain any of the underlying physical characteristics of the brightness simulation, which itself is quite inaccurate. Its vibes all the way down
This is all fine if you’re doing visualisation, but this isn’t an accurate simulation. I wish this was just called a visualisation of a black hole, but its being communicated as if this is super hard science with credentials and all
20k
如果有人关心准确性,这只是一个“可视化级别”的软件。让我有点恼火的是,人们把这些模拟描述成物理上非常精确,但实际上它们存在重大不准确之处,有些非常明显甚至是刻意的。这个模拟有一些严重的物理限制。
如果它没有特别标榜自己是“物理精确的”黑洞模拟,我完全不会介意,但这就属于具有误导性的科学传播,而这种问题常常被轻描淡写地忽略,仿佛无关紧要,所以我们得澄清一些事情!
-
吸积盘不应该是红色的,人们只是觉得红色好看才期待它这样。黑洞吸积盘几乎普遍热到呈现蓝色。《星际穿越》也这么做了,而且试图用非常没有说服力的理由搪塞过去。
-
这是一个非自转/低自转(?)的黑洞,这并不非常现实。
-
它忽略了摄像机的位置(这会影响洛伦兹偏移)。
-
多普勒偏移并不非常准确。
-
它没有以任何精度模拟吸积盘的温度分布或颜色。通常你会把吸积盘建模为黑体辐射体,通过多普勒效应进行偏移,然后与人类视觉响应(LMS)卷积,转换到XYZ再到RGB,进行物理色调映射,最后进行sRGB转换。而这个模拟完全没做这些——没有任何一步是物理精确的。它甚至在示意层面上都不正确,我们稍后会谈到。
-
实际的辐射传输与真实使用的模型相比简化了很多,并基于任何实际数据,只用了非常简化的方程。盘的光学厚度和发射率是任意的,尺寸也是任意的,没有正确纳入亮度或消光。例如这里(https://github.com/aplavin/blackhole.plav.in/blob/2f004bfeca670d4dfdaf24dc5c34af862f5f867b/src/raymarch.frag#L379)是极度简化的。
-
多普勒计算使用了错误的方程。他们用的是I^3变体,而从盘样本中得到的其实是“辐射通量”,实际应为F_obs = F_emit / (z+1)^4。这是图像处理中非常常见的错误,导致多普勒偏移和观测者亮度不正确。在这里(https://github.com/aplavin/blackhole.plav.in/blob/2f004bfeca670d4dfdaf24dc5c34af862f5f867b/src/raymarch.frag#L47)的“表面亮度”其实不是光谱辐射度,而是辐射通量。
像亮度到颜色的转换尤其不准确。例如,如果你查看源码: https://github.com/aplavin/blackhole.plav.in/blob/2f004bfeca670d4dfdaf24dc5c34af862f5f867b/src/raymarch.frag#L140 它完全随意地将伪亮度映射到颜色。由此产生的颜色/亮度与任何物理量都不对应。它还进行了从线性量(亮度)到sRGB的“线性”映射(而sRGB是非线性过程!!),这意味着它甚至没有保留亮度模拟的任何底层物理特征——而亮度模拟本身就已经很不准确了。从头到尾都只是凭感觉。
如果你只是在做可视化,这完全没问题,但这不是一个精确的模拟。我希望它只是被称作黑洞的可视化,但它却被宣传成有资质的、极其严谨的科学成果。
AI companies are shredding rare books #
https://news.ycombinator.com/item?id=49069602
I’ve limited sympathy for the publishers.
It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There’s no real need for any of these so-called rare books to be rare while they’re under copyright.
And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are still in good nick. It’s ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.
If we have to have copyright laws, I’d like to see two changes to them.
When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print, without compensation to the original publisher, and with renegotiated royalties for the author.
And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations’ time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.
squidbeak
我对出版商没什么同情。
想到他们能把作品一直攥在手里直到版权过期,让它们绝版,就让我火大。在这些作品受版权保护期间,所谓的"稀有书籍"根本没必要稀有。
与此相关的是,那些仍在印行的书,大多也是以最糟糕的方式在印。我经常看到17或18世纪装帧精良的书籍如今品相依然完好。荒谬的是,到了21世纪,出版标准竟沦落到大多数作品只有一次性装订版本可选——哪怕花再多钱也买不到一本真正像样的精装本。
如果非要保留版权法,我希望看到两点修改:
当出版商没有动力继续印刷某个版本时,应允许其他出版商印刷,无需向原出版商支付补偿,并与作者重新协商版税。
如果出版商将一本书保持在印——但只使用畅销书级别的材料:遇潮就起皱的卫生纸级纸张、几个旱季之后就会掉页的胶装——并且拒绝印刷带有书帖、优质纸张和精美印刷的耐久精装本(那种几代后仍可阅读的版本)——那么任何有意尝试的其他出版商都应该能够参与竞争,同样无需向原出版商支付补偿,不过在这种情况下,或许需要向作者支付相应的版税。
How is the Bun rewrite in Rust going? #
https://news.ycombinator.com/item?id=49068644
I really do not understand how software developer think anymore. Using a LLM to translate a project in a short time, is by itself incredible. Just like one-shot whatever office clone.
But what makes software is not the fast creation of a “product” but that actual development of its features. Figuring out how everything needs to work together, fixing the bugs, and the o so boring UI work.
I have used LLMs to create stuff like word clones just for fun. It was a disaster. Sure, it had the basic functionality. But the moment you started with page structure (harder then one non-stop scrolling page), tables, images, rotating, and so many details that make up just the basics of word. Not even the extended functionality. You see every LLM just fall on its face.
Sure, i can clone sqlite from c to rust. Hell, i may even get it to do all the tests 100%. But there is a 99% chance that the clone will be slower, as it lacks the years of optimizations from the original language. There will be new bugs because of the language changeover. There is a need for future support and fixes.
People threat software like its something it is not. But unlike the past where your clients question your sanity for charging 100k for a piece of software. Not understanding its not just about writing the code. Now those expectation are even more pushed forwards, because of articles like this.
I constantly see software being published on reddit that does X, Y, Z only for the authors to abandon it as fast as they vibe coded it. Because fixing bugs is NOT sexy. Even with a LLM at your fingertips. Dealing with nagging users, is not sexy. Dealing with security issues, is NOT sexy. Dealing with data structure / databases, especially as your system changes … you get the point.
Not understanding to the core the software you wrote, is going to exploded in your face.
This is why these stupid “we rewrote X into Z with a LLM in Y days” mean nothing. Its one thing to get a head start using this trick, its another to actually learn the code of your rewrite. And dedicated the time into maintaining the port, growing it, fixing it. This is where a lot of software fails. But now this crap is out there, instead of the maintained version of Zig, now we have a unmaintained Rust version that clouded the airwaves because if people now search for it, those articles “X in Z days” will pop up.
What have we become …
benjiro29
我真的越来越搞不懂软件开发者的思维方式了。用大语言模型在短时间内翻译一个项目,这本身确实不可思议——就像一键生成一个办公软件克隆版。
但软件的价值并不在于快速制造一个"产品”,而在于其功能真正的开发过程:弄清楚所有部分如何协同工作、修复漏洞,以及那极其枯燥的界面工作。
我也曾用大语言模型试着做过文字处理软件克隆版,纯粹为了好玩。结果简直是一场灾难。它确实具备基础功能,但一旦涉及页面结构(比连续滚动的页面难得多)、表格、图片、旋转功能,以及构成Word基础功能的各种细节(更别说那些进阶功能了),大语言模型立刻就崩溃了。
当然,我可以用大语言模型把SQLite从C语言克隆到Rust语言,甚至可能让所有测试用例100%通过。但有99%的概率这个克隆版会比原版更慢——因为它缺少原版语言积累多年的优化。语言转换还会带来新漏洞,而未来的维护与修复工作更是必不可少。
人们总把软件当成并非其本质的东西。过去客户质疑你收10万美元开发一套软件是否疯了,他们不理解这不只是写代码。而现在,因为这类文章,这种期望值被推得更高了。
我经常看到Reddit上有人发布实现X、Y、Z功能的软件,结果作者们像写"氛围代码"一样快速抛弃它。因为修bug一点也不酷——就算有大语言模型在手也一样。应付烦人的用户不酷,处理安全问题不酷,处理数据结构/数据库(尤其是当系统不断变化时)……你懂的。
如果不深入理解自己写的软件,迟早会翻车。
这就是为什么那些愚蠢的"我们用大语言模型在Y天内把X重写成了Z"毫无意义。用这个技巧抢占先机是一回事,真正学习你重写的代码并投入时间维护、改进、修复它,则是另一回事。这正是很多软件失败的原因。但现在这种垃圾充斥市场:原本维护良好的Zig版本,却被一个无人维护的Rust版本污染了信息环境——因为人们搜索时,那些"X在Z天内完成"的文章就会跳出来。
我们到底变成了什么样……
Kimi-K3 on HuggingFace #
https://news.ycombinator.com/item?id=49066701
Even if the output is like 5-6 tok/s, that might be usable for some purposes.
You’ll spend ~100x more on electricity than the API cost to have it run on someone else’s GPU at several hundred tokens per second.
I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really really narrow. I wouldn’t be surprised if this is an empty set.
fooker
即使输出速度只有5-6 tok/s,对某些用途来说或许也能用。你在电费上的开销大约会是使用他人GPU以每秒数百token运行API成本的100倍。我认为只有极端的数据隐私需求才能证明这种做法的合理性,但{需要绝对数据隐私、需要运行最先进模型、买不起GPU}这三者的交集真的非常非常狭窄。如果这是个空集,我也不会感到意外。