Skip to content

周报 2026-08-03 ~ 2026-08-09

生成时间:2026/8/9 11:02:08(UTC: 2026-08-09T03:02:08.650Z)

本周自动总结未启用或调用失败,以下为原始内容合并。

内部采集快照,不作为站点文章发布。

生成时间:2026/8/3 09:30:03(UTC: 2026-08-03T01:30:03.456Z)

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

Section titled “AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis”

👍 292 · arXiv

Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists. As a resul…

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Section titled “Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents”

👍 285 · arXiv

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execu…

👍 257 · arXiv

Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However…

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Section titled “Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering”

👍 168 · arXiv

Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this c…

PhiZero: A World Model Built Around Physical Language

Section titled “PhiZero: A World Model Built Around Physical Language”

👍 159 · arXiv

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future video…

  • State safety and recovery: protect persisted data with a quarantine store that survives primary-database damage, crash-recoverable SQLite snapshots, crash-durable fi…

链接https://github.com/openclaw/openclaw/releases/tag/v2026.7.2-beta.7

On the latest episode of Equity, we discuss why Sam Altman has calling on the industry to “pace the rate of AI development.”

来源TechCrunch AI

Judge denies xAI’s request to block Minnesota ban on ‘nudify’ apps

Section titled “Judge denies xAI’s request to block Minnesota ban on ‘nudify’ apps”

Despite a lawsuit from xAI, a Minnesota ban on apps that allow users to “nudify” images can move forward.

来源TechCrunch AI

YouTuber Hank Green says his AI usage is ‘not healthy’

Section titled “YouTuber Hank Green says his AI usage is ‘not healthy’”

Green offered a remarkable apology, saying that “the level of dopamine that I’ve been getting from interacting with LLMs … is not healthy for me or good for the world.”

来源TechCrunch AI

Sam Altman is still making the case for parenting via ChatGPT

Section titled “Sam Altman is still making the case for parenting via ChatGPT”

OpenAI’s CEO seemed excited to share a “cool use case” for parents.

来源TechCrunch AI

This $9 key physically locks your most addictive apps

Section titled “This $9 key physically locks your most addictive apps”

This $9 NFC key requires you to physically scan it to unlock distracting apps on your phone.

来源TechCrunch AI

OpenAI reportedly finds evidence that more of its agents ran amok

Section titled “OpenAI reportedly finds evidence that more of its agents ran amok”

OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.

来源TechCrunch AI

India is starting to pay for apps, not just download them

Section titled “India is starting to pay for apps, not just download them”

India’s app market generated a record $345 million in Q2.

来源TechCrunch AI

Google nixes its Earth AI feature one day after launch, amid criticism it would spread misinformation

Section titled “Google nixes its Earth AI feature one day after launch, amid criticism it would spread misinformation”

A tool that allowed anyone to generate fake AI-generated imagery and superimpose it over real Google Earth maps quickly spurred backlash.

来源TechCrunch AI


内部采集快照,不作为站点文章发布。

生成时间:2026/8/4 09:19:37(UTC: 2026-08-04T01:19:37.142Z)

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Section titled “Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents”

👍 294 · arXiv

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execu…

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Section titled “From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement”

👍 74 · arXiv

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability rem…

👍 53 · arXiv

World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is d…

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Section titled “Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory”

👍 51 · arXiv

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-…

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

Section titled “N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens”

👍 50 · arXiv

We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offlin…

  • npm plugin updates: accept singleton-array metadata from newer npm clients so tracked official plugins can install and update to correction releases. (#108336)

链接https://github.com/openclaw/openclaw/releases/tag/v2026.7.1-2

Release 0.147.0-alpha.6

链接https://github.com/openai/codex/releases/tag/rust-v0.147.0-alpha.6

AI’s debt binge can’t last, hidden borrowing reaches $1.65T

Section titled “AI’s debt binge can’t last, hidden borrowing reaches $1.65T”

Article URL: https://fortune.com/2026/07/31/ai-debt-hypescalers-capex-capital-spending-hidden-borrowing-bond-issuance/ Comments URL: https://news.ycombinator.com/item?id=49160699 Points: 112

来源Hacker News AI

What’s the largest software project AI can complete on its own?

Section titled “What’s the largest software project AI can complete on its own?”

Article URL: https://epoch.ai/MirrorCode Comments URL: https://news.ycombinator.com/item?id=49157786 Points: 66

来源Hacker News AI

The AI bubble is popping; we just don’t know it yet

Section titled “The AI bubble is popping; we just don’t know it yet”

Article URL: https://www.theregister.com/ai-and-ml/2026/08/03/the-ai-bubble-is-already-popping-we-just-dont-know-it-yet/5282004 Comments URL: https://news.ycombinator.com/item?id=49154601 Points: 75

来源Hacker News AI

Article URL: https://research.jfrog.com/post/sqlite-critical-cves-or-llm-slops/ Comments URL: https://news.ycombinator.com/item?id=49154332 Points: 698

来源Hacker News AI

Show HN: Nightcrawler – A local AI pentesting agent running on a smartphone

Section titled “Show HN: Nightcrawler – A local AI pentesting agent running on a smartphone”

Article URL: https://github.com/garagehq/nightcrawler/ Comments URL: https://news.ycombinator.com/item?id=49154127 Points: 102

来源Hacker News AI

Prevent cognitive debt by manually retyping LLM-generated code

Section titled “Prevent cognitive debt by manually retyping LLM-generated code”

Article URL: https://ankursethi.com/blog/prevent-cognitive-debt-by-manually-retyping-llm-generated-code/ Comments URL: https://news.ycombinator.com/item?id=49153374 Points: 380

来源Hacker News AI

Article URL: https://bjorg.bjornroche.com/management/ai-productivity-gap/ Comments URL: https://news.ycombinator.com/item?id=49152222 Points: 105

来源Hacker News AI

AI migrated legacy COBOL programs to Java, bugs included

Section titled “AI migrated legacy COBOL programs to Java, bugs included”

Article URL: https://arxiv.org/abs/2607.28271 Comments URL: https://news.ycombinator.com/item?id=49150773 Points: 87

来源Hacker News AI


内部采集快照,不作为站点文章发布。

生成时间:2026/8/5 09:23:42(UTC: 2026-08-05T01:23:42.230Z)

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Section titled “SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks”

👍 140 · arXiv

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices …

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Section titled “LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks”

👍 127 · arXiv

Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses…

👍 56 · arXiv

On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged information but can induce a privilege illusion: the student l…

👍 50 · arXiv

On-policy distillation (OPD), which aligns a student with the teacher’s token-level distribution on the student’s own rollouts, is an effective paradigm for transferring capabilities across LLMs. Prev…

Progressive Agent Skill Generation via Reinforcement Learning

Section titled “Progressive Agent Skill Generation via Reinforcement Learning”

👍 49 · arXiv

Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially designed for different evidence sources. In contrast, learning-based approaches of…

  • npm plugin updates: accept singleton-array metadata from newer npm clients so tracked official plugins can install and update to correction releases. (#108336)

链接https://github.com/openclaw/openclaw/releases/tag/v2026.7.1-2

  • Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model’s MTP head for speculative decoding automatically
  • /v1/chat/completions streaming now matches OpenAI’s w…

链接https://github.com/ollama/ollama/releases/tag/v0.32.6-rc0

Release 0.147.0-alpha.7

链接https://github.com/openai/codex/releases/tag/rust-v0.147.0-alpha.7

AI fuels more than half of cybercrime in Africa as scams surge – Interpol

Section titled “AI fuels more than half of cybercrime in Africa as scams surge – Interpol”

https://www.interpol.int/Media/Documents/Publications/Cyberc

Comments URL: https://news.ycombinator.com/item?id=49175826 Points: 129

来源Hacker News AI

Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

Section titled “Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]”

Article URL: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf Comments URL: https://news.ycombinator.com/item?id=49175717 Points: 56

来源Hacker News AI

AI Data Centers Are Driving Up Power Bills – This Map Shows Where

Section titled “AI Data Centers Are Driving Up Power Bills – This Map Shows Where”

Article URL: https://www.gadgetreview.com/ai-data-centers-are-driving-up-power-bills-this-map-shows-where Comments URL: https://news.ycombinator.com/item?id=49172433 Points: 61

来源Hacker News AI

Article URL: https://www.warp.dev/blog/introducing-the-warp-agent-cli-coding-agent Comments URL: https://news.ycombinator.com/item?id=49171766 Points: 93

来源Hacker News AI

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Section titled “When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation”

Article URL: https://arxiv.org/abs/2602.16763 Comments URL: https://news.ycombinator.com/item?id=49170915 Points: 78

来源Hacker News AI

Article URL: https://www.wheresyoured.at/the-ai-demand-bubble/ Comments URL: https://news.ycombinator.com/item?id=49170648 Points: 104

来源Hacker News AI

Agent skills that bring team coding standards to Claude Code and Codex

Section titled “Agent skills that bring team coding standards to Claude Code and Codex”

Article URL: https://github.com/tikalk/adlc-team-skills Comments URL: https://news.ycombinator.com/item?id=49169640 Points: 74

来源Hacker News AI

It’s not a fear of “AI communism”; it’s a fear of competitive market capitalism

Section titled “It’s not a fear of “AI communism”; it’s a fear of competitive market capitalism”

Article URL: http://observationalepidemiology.blogspot.com/2026/07/its-not-fear-of-ai-communism-its-fear.html Comments URL: https://news.ycombinator.com/item?id=49169227 Points: 78

来源Hacker News AI


内部采集快照,不作为站点文章发布。

生成时间:2026/8/6 09:21:04(UTC: 2026-08-06T01:21:04.906Z)

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Section titled “MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations”

👍 87 · arXiv

Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-T…

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Section titled “JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion”

👍 79 · arXiv

Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a …

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

Section titled “AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling”

👍 74 · arXiv

Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. …

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

Section titled “Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing”

👍 71 · arXiv

Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained…

InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

Section titled “InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis”

👍 56 · arXiv

Single-image feed-forward 3D Gaussian Splatting (3DGS) aims to directly generate a renderable 3D scene representation from one input image, avoiding the cost of multi-view capture and per-scene optimi…

Changes since langchain-anthropic==1.5.3

release(anthropic): 1.5.4 (#39277) fix(anthropic): handle tool schemas with unsupported top-level composition (#39273) chore: bump the minor-and-patch group a…

链接https://github.com/langchain-ai/langchain/releases/tag/langchain-anthropic%3D%3D1.5.4

  • Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model’s MTP head for speculative decoding automatically
  • /v1/chat/completions streaming now matches OpenAI’s w…

链接https://github.com/ollama/ollama/releases/tag/v0.32.6

  • Bump Flow canary on release
  • Add URLReadTool for reading arbitrary URLs
  • Add app metadata to platform action tools
  • Unify scaffolding under `crewai create <resourc…

链接https://github.com/crewAIInc/crewAI/releases/tag/1.15.12

Release 0.147.0-alpha.13

链接https://github.com/openai/codex/releases/tag/rust-v0.147.0-alpha.13

Article URL: https://www.primeintellect.ai/blog/prime-agent Comments URL: https://news.ycombinator.com/item?id=49189075 Points: 95

来源Hacker News AI

Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery

Section titled “Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery”

Article URL: https://www.wired.com/story/meta-ran-ads-that-contained-ai-generated-child-sexual-abuse-imagery/ Comments URL: https://news.ycombinator.com/item?id=49187977 Points: 244

来源Hacker News AI

Born Against, or why hobby programming communities are against LLM usage

Section titled “Born Against, or why hobby programming communities are against LLM usage”

Article URL: https://blog.fogus.me/llm/born-against.html Comments URL: https://news.ycombinator.com/item?id=49187061 Points: 123

来源Hacker News AI

Microsoft’s AI Sales Mostly Come from OpenAI, Disclosures Show

Section titled “Microsoft’s AI Sales Mostly Come from OpenAI, Disclosures Show”

https://www.bloomberg.com/news/articles/2026-08-05/microsoft

Comments URL: https://news.ycombinator.com/item?id=49186766 Points: 61

来源Hacker News AI

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

Section titled “Beating GPT-5.6 Sol on retrieval with 100x cheaper open models”

Article URL: https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency Comments URL: https://news.ycombinator.com/item?id=49186762 Points: 219

来源Hacker News AI

Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025)

Section titled “Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (2025)”

Article URL: https://arxiv.org/abs/2510.01395 Comments URL: https://news.ycombinator.com/item?id=49186720 Points: 70

来源Hacker News AI

TIME Is Serving AI Bots a Different Website, with Ads Built In

Section titled “TIME Is Serving AI Bots a Different Website, with Ads Built In”

Article URL: https://www.vincentschmalbach.com/time-serves-ai-bots-a-different-website/ Comments URL: https://news.ycombinator.com/item?id=49182041 Points: 230

来源Hacker News AI

Anthropic AI created fake profiles and impersonated people in attempted hack

Section titled “Anthropic AI created fake profiles and impersonated people in attempted hack”

Article URL: https://www.bbc.co.uk/news/articles/c1w1lvn7d9go Comments URL: https://news.ycombinator.com/item?id=49181773 Points: 50

来源Hacker News AI


内部采集快照,不作为站点文章发布。

生成时间:2026/8/7 10:01:00(UTC: 2026-08-07T02:01:00.211Z)

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

Section titled “ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment”

👍 55 · arXiv

Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agent…

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Section titled “ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation”

👍 46 · arXiv

Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of…

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

Section titled “Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes”

👍 42 · arXiv

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms o…

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

Section titled “The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads”

👍 37 · arXiv

Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user…

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

Section titled “OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents”

👍 28 · arXiv

LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goal…

Changes since langchain-anthropic==1.5.3

release(anthropic): 1.5.4 (#39277) fix(anthropic): handle tool schemas with unsupported top-level composition (#39273) chore: bump the minor-and-patch group a…

链接https://github.com/langchain-ai/langchain/releases/tag/langchain-anthropic%3D%3D1.5.4

  • Bump Flow canary on release
  • Add URLReadTool for reading arbitrary URLs
  • Add app metadata to platform action tools
  • Unify scaffolding under `crewai create <resourc…

链接https://github.com/crewAIInc/crewAI/releases/tag/1.15.12

  • Install portable Agent Plugins and search across local, personal, workspace, and remote plugin catalogs. (#36544, #36409, #36919, #36796)
  • Organize conversations into persistent, ma…

链接https://github.com/openai/codex/releases/tag/rust-v0.147.0

OpenAI’s new AI smart speaker will reportedly sell for between $300 and $400

Section titled “OpenAI’s new AI smart speaker will reportedly sell for between $300 and $400”

Additional details about OpenAI’s mysterious new AI device make it sound like a pricey smart speaker.

来源TechCrunch AI

ChatGPT brings unlimited text chats to free users

Section titled “ChatGPT brings unlimited text chats to free users”

OpenAI said that ChatGPT free and Go users are also getting a new think button for complex queries.

来源TechCrunch AI

Naïve raises $28.5M to automate the grunt work of setting up and running a company

Section titled “Naïve raises $28.5M to automate the grunt work of setting up and running a company”

Taking vibe-coding a step further, Naïve claims its infra can automate most of the work in setting up and running a business.

来源TechCrunch AI

Gen Z dating apps like Ditto ditch swiping in favor of AI matchmaking

Section titled “Gen Z dating apps like Ditto ditch swiping in favor of AI matchmaking”

This generation of twentysomethings is so disillusioned with swipe-based dating apps that they’ll try literally anything else — even an AI matchmaker.

来源TechCrunch AI

OpenAI says Apple’s own security practices undermine its trade secrets case

Section titled “OpenAI says Apple’s own security practices undermine its trade secrets case”

Newly filed court exhibits show OpenAI’s legal strategy in Apple’s trade secrets lawsuit: argue that Apple’s own security and offboarding practices — including allowing an Apple manager to access a former engineer’s iCloud account after he left the company —undermine its claims that the allegedly st

来源TechCrunch AI

Section titled “Amid legal battles, Suno says it will start watermarking songs”

Suno’s watermarking feature comes as the company is fighting legal battles on several fronts.

来源TechCrunch AI

Ex-Spotify employees raise $10M to bring the AI behind its recommendations to e-commerce

Section titled “Ex-Spotify employees raise $10M to bring the AI behind its recommendations to e-commerce”

The startup’s platform predicts which product a shopper wants next, learns their general taste, and fine-tunes continuously based on what they do in real time.

来源TechCrunch AI

Exclusive: Mirendil inks $100M+ Google Cloud deal to scale self-improving AI

Section titled “Exclusive: Mirendil inks $100M+ Google Cloud deal to scale self-improving AI”

Mirendil has signed a $100 million-plus Google Cloud partnership to expand its compute infrastructure, powering research into self-improving AI systems designed to accelerate scientific discovery and AI development.

来源TechCrunch AI


内部采集快照,不作为站点文章发布。

生成时间:2026/8/8 08:44:41(UTC: 2026-08-08T00:44:41.845Z)

Recursive Synthesis for Long-Horizon Terminal Tasks

Section titled “Recursive Synthesis for Long-Horizon Terminal Tasks”

👍 212 · arXiv

High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, …

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Section titled “AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning”

👍 67 · arXiv

Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, mul…

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

Section titled “ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment”

👍 61 · arXiv

Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agent…

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Section titled “OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models”

👍 52 · arXiv

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent’s actions, states, and reasoning. Verifying whether it fulfilled the task instruction is…

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

Section titled “Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval”

👍 46 · arXiv

Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings…

Changes since langchain-openai==1.4.1

release(openai): 1.4.2 (#39322) fix(openai): handle ContextWindowExceededError (#39300) chore: bump the minor-and-patch group across 3 directories with 7 updat…

链接https://github.com/langchain-ai/langchain/releases/tag/langchain-openai%3D%3D1.4.2

  • Fix preservation of provider on LiteLLM-routed models.
  • Harden brittle LLM event-bus mocks.
  • Fix underreporting of Anthropic cache token usage.
  • Bump h2 to versio…

链接https://github.com/crewAIInc/crewAI/releases/tag/1.15.13

  • Install portable Agent Plugins and search across local, personal, workspace, and remote plugin catalogs. (#36544, #36409, #36919, #36796)
  • Organize conversations into persistent, ma…

链接https://github.com/openai/codex/releases/tag/rust-v0.147.0

Article URL: https://www.databricks.com/blog/managing-ai-coding-costs-scale Comments URL: https://news.ycombinator.com/item?id=49214468 Points: 155

来源Hacker News AI

Oracle bans AI-generated code from OpenJDK

Section titled “Oracle bans AI-generated code from OpenJDK”

Article URL: https://app.dealroom.co/news/feed/oracle-bans-ai-generated-code-from-openjdk-despite-ellison-s-claim-oracle-isn-t-writing-its-own-code Comments URL: https://news.ycombinator.com/item?id=49213754 Points: 371

来源Hacker News AI

AI psychosis is the new leadership blind spot

Section titled “AI psychosis is the new leadership blind spot”

Article URL: https://www.fastcompany.com/91576086/ai-psychosis-is-the-new-leadership-blind-spot-ai-leadership-blind-spots Comments URL: https://news.ycombinator.com/item?id=49210077 Points: 159

来源Hacker News AI

Kitesurf: Agent-first browser that runs in V8 isolates

Section titled “Kitesurf: Agent-first browser that runs in V8 isolates”

Article URL: https://blog.cloudflare.com/kitesurf/ Comments URL: https://news.ycombinator.com/item?id=49208393 Points: 160

来源Hacker News AI

Article URL: https://mccormick.cx/news/entries/why-i-won-t-read-llm-authored-fiction Comments URL: https://news.ycombinator.com/item?id=49207146 Points: 70

来源Hacker News AI

New Orleans is testing Carbyne’s AI-powered Emergency Call Triage software

Section titled “New Orleans is testing Carbyne’s AI-powered Emergency Call Triage software”

Article URL: https://www.shreveporttimes.com/story/news/local/louisiana/2026/07/28/is-new-orleans-using-ai-to-answer-911-calls-instead-of-human-dispatchers-impacts-emergencies-crime/91065014007/ Comments URL: https://news.ycombinator.com/item?id=49204546 Points: 72

来源Hacker News AI

Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)

Section titled “Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)”

Article URL: https://www.aleksagordic.com/blog/vllm Comments URL: https://news.ycombinator.com/item?id=49202852 Points: 142

来源Hacker News AI

Article URL: https://illegal.solutions/posts/xai_pollution Comments URL: https://news.ycombinator.com/item?id=49201342 Points: 144

来源Hacker News AI