TUJLTECH UPDATES, JUST LATEST

DAILY AI BRIEFING · 每日 AI 简报

让变化经过筛选,
再抵达你。

科技更新,只看最新。每天从可靠来源中挑选值得留下的 AI 动向;中英对照,链接回到原文,事实与判断分开。

READ THIS EDITION ↓

THE SIGNAL / 今日信号

六件值得知道的事

6 stories · 约 8 分钟

01

TASTE 基准测试模型判断 AI 安全研究提案的能力

TASTE tests whether models can judge AI-safety research proposals

Anthropic Fellows 发布 TASTE,以 92 组两两比较测试模型能否判断 AI 安全研究提案。团队估算人类研究者的一致率为 77%,表现最佳的 Fable 5 为 60%。这是规模较小的早期基准,置信区间较宽;结果表明模型仍落后于人类判断,而不是证明某个系统具备可靠的研究评审能力。

Anthropic Fellows released TASTE, a benchmark of 92 pairwise comparisons asking models to judge AI-safety research proposals. The team estimates human researcher agreement at 77%, while the best model, Fable 5, reached 60%. This is a small early benchmark with wide confidence intervals; it suggests a gap from human judgment rather than establishing reliable automated peer review.

阅读原始来源 / Read original source
02

SageMaker Feature Store 新增批量写入与记录发现 API

SageMaker Feature Store adds batch-write and record-discovery APIs

AWS 为 SageMaker Feature Store 增加 BatchWriteRecord 与 ListRecords。前者单次调用最多可向多个特征组写入 25 条记录,支持部分成功、单条 TTL 与 EventTime 顺序;后者可枚举 Standard 和 In-Memory 层的记录 ID。这是官方产品更新,实际吞吐与成本仍取决于工作负载。

AWS has added BatchWriteRecord and ListRecords to SageMaker Feature Store. One call can write up to 25 records across multiple feature groups, with partial success, per-record TTL, and EventTime ordering; ListRecords enumerates record IDs in the Standard and In-Memory tiers. This is an official product update, while real throughput and cost will depend on the workload.

阅读原始来源 / Read original source
03

Decathlon 分享用 Chronos-2 规模化运行需求预测的流程

Decathlon details its at-scale demand-forecasting workflow with Chronos-2

Decathlon 与 AWS 介绍其 Chronos-2 需求预测流程:每个区域最多覆盖 25,000 件商品,同时生成 12 周和 52 周预测;模型约每六个月微调,预测每周批量运行。这是单一企业与供应商联合案例,说明一种工程路径,不代表其他零售场景会得到相同结果。

Decathlon and AWS describe a Chronos-2 demand-forecasting workflow covering up to 25,000 products per zone with 12- and 52-week horizons. Models are fine-tuned roughly every six months and forecasts run in weekly batches. This is a joint customer-and-vendor case that illustrates one engineering path, not a promise of identical results elsewhere.

阅读原始来源 / Read original source
04

SageMaker Inference Components 增加跨可用区平衡策略

SageMaker Inference Components adds availability-zone balancing

AWS 为 CreateInferenceComponent 新增 SchedulingConfig,可用 AvailabilityZoneBalance 在 SPREAD 与 BINPACK 之间选择,帮助团队在高可用与资源密度之间取舍。Salesforce 案例称通过共置模型将基础设施成本降低 8 倍;这一数字来自客户案例,应按其特定架构与流量条件理解。

AWS has added SchedulingConfig to CreateInferenceComponent, using AvailabilityZoneBalance to choose between SPREAD and BINPACK placement for different reliability and density goals. A Salesforce case reports an 8× infrastructure-cost reduction from model co-hosting; that figure belongs to the customer's specific architecture and traffic conditions.

阅读原始来源 / Read original source
05

Gemini Notebook 改用五小时刷新一次的弹性额度

Gemini Notebook moves to flexible limits that refresh every five hours

Google Labs 调整 Gemini Notebook 的使用额度:计算相关上限从每日刷新改为每五小时刷新,并会依据提示复杂度、对话长度、来源数量与功能而变化;视频概览和幻灯片等任务可排队延后生成并发送通知。更新计划 9 月 2 日开始推出,属于产品公告。

Google Labs is changing Gemini Notebook's compute limits from daily resets to refreshes every five hours, with consumption varying by prompt complexity, chat length, source count, and feature. Video Overviews and Slide Decks can be deferred and notify users when ready. The product change is scheduled to begin rolling out on September 2.

阅读原始来源 / Read original source
06

AWS 展示把持续现代化嵌入 CI/CD 的自建方案

AWS demonstrates a do-it-yourself continuous-modernization pipeline

AWS 发布一套用 AWS Transform custom、GitHub Actions 与 Dependabot 构建持续现代化流水线的技术方案,示例覆盖依赖修复、自动文档、跨仓库变换与知识复用。文章明确示例在非交互模式中使用信任全部工具的参数,并提醒生产环境先审查安全策略;这是技术演示,不是无人监督变更的普遍背书。

AWS has published a technical pattern combining AWS Transform custom, GitHub Actions, and Dependabot for dependency remediation, automatic documentation, cross-repository transformations, and knowledge reuse. The walkthrough explicitly uses a trust-all-tools flag for non-interactive runs and advises reviewing security policy before production use. It is a demonstration, not a general endorsement of unsupervised changes.

阅读原始来源 / Read original source

EDITOR'S NOTE · 编者的话

今天的共同线索,是让 AI 经得起日常运行。

批量写入、跨可用区部署、资源额度、人工判断和持续维护看似属于不同层次,却都在回答同一个问题:能力如何稳定地进入真实系统。事实是这些 API、产品调整、架构案例和研究结果已经发布;我们的判断是,AI 的下一段进步不仅来自模型更强,也来自边界更清楚、流程更可检查。案例数字与早期研究结论均按发布方报告理解。

Batch writes, multi-zone deployment, usage budgets, human judgment, and continuous maintenance may sit at different layers, but they answer the same question: how does capability operate reliably inside real systems? The fact is that these APIs, product changes, architecture cases, and research results were published. Our judgment is that the next stage of AI progress will come not only from stronger models, but also from clearer boundaries and more inspectable processes. Case figures and early research findings remain publisher-reported.

TWO ADDRESSES · 两个地址

一个保存我们,
一个观察世界。

RUJF.AIRecord Us Just Forever

保存属于我们、记忆与创作的东西。

TUJL.COMTech Updates, Just Latest

观察外面的世界,记录 AI 此刻正在发生什么。

站外推广狗狗加速