<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Alien Life AI: companies and labs</title>
    <link>https://alienlifeai.com/companies.html</link>
    <atom:link href="https://alienlifeai.com/companies.xml" rel="self" type="application/rss+xml"/>
    <description>Primary-source posts from AI labs and companies: their own newsrooms, research blogs and papers. Every item links to the publisher; this feed republishes nobody's work.</description>
    <language>en</language>
    <lastBuildDate>Tue, 22 Sep 2026 23:12:20 +0000</lastBuildDate>
    <item>
      <title>Better prompt caching for GPT-6</title>
      <link>https://openai.com/index/better-prompt-caching-for-gpt-6</link>
      <guid isPermaLink="true">https://openai.com/index/better-prompt-caching-for-gpt-6</guid>
      <pubDate>Tue, 22 Sep 2026 21:00:00 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">OpenAI</source>
      <description>Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.</description>
    </item>
    <item>
      <title>What’s New for Game Developers: DLSS 5 with 3D-Guided Neural Rendering, NVIDIA ACE Updates, and New RTX Kit Capabilities</title>
      <link>https://developer.nvidia.com/blog/whats-new-for-game-developers-dlss-5-with-3d-guided-neural-rendering-nvidia-ace-updates-and-new-rtx-kit-capabilities/</link>
      <guid isPermaLink="true">https://developer.nvidia.com/blog/whats-new-for-game-developers-dlss-5-with-3d-guided-neural-rendering-nvidia-ace-updates-and-new-rtx-kit-capabilities/</guid>
      <pubDate>Tue, 22 Sep 2026 20:48:36 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">NVIDIA</source>
      <description>NVIDIA DLSS 5 introduces DLSS 3D-Guided Neural Rendering and granular controls that help game developers add lifelike lighting and material detail while...</description>
    </item>
    <item>
      <title>Bring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock</title>
      <link>https://aws.amazon.com/blogs/machine-learning/bring-more-intelligence-to-everyday-work-with-gpt-6-sol-and-gpt-6-luna-on-amazon-bedrock/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/bring-more-intelligence-to-everyday-work-with-gpt-6-sol-and-gpt-6-luna-on-amazon-bedrock/</guid>
      <pubDate>Tue, 22 Sep 2026 18:10:22 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>GPT-6 Sol and GPT-6 Luna are now generally available on Amazon Bedrock, giving you more options to match intelligence and efficiency to each workload.</description>
    </item>
    <item>
      <title>Introducing GPT-6 Sol and Luna</title>
      <link>https://openai.com/index/introducing-gpt-6-sol-and-luna</link>
      <guid isPermaLink="true">https://openai.com/index/introducing-gpt-6-sol-and-luna</guid>
      <pubDate>Tue, 22 Sep 2026 18:00:00 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">OpenAI</source>
      <description>Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.</description>
    </item>
    <item>
      <title>Claude Opus 5.5 is now available on AWS</title>
      <link>https://aws.amazon.com/blogs/machine-learning/claude-opus-5-5-is-now-available-on-aws/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/claude-opus-5-5-is-now-available-on-aws/</guid>
      <pubDate>Tue, 22 Sep 2026 17:28:01 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>Claude Opus 5.5, Anthropic&#x27;s most capable Opus model for agentic coding, knowledge work, and long-running tasks, is now available on Amazon Bedrock and Claude Platform on AWS. This post covers what&#x27;s new in Opus 5.5, practical guidance, and how to start building with the model on Amazon Bedrock.</description>
    </item>
    <item>
      <title>Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing</title>
      <link>https://developer.nvidia.com/blog/enabling-private-high-performance-production-ai-inference-with-nvidia-confidential-computing/</link>
      <guid isPermaLink="true">https://developer.nvidia.com/blog/enabling-private-high-performance-production-ai-inference-with-nvidia-confidential-computing/</guid>
      <pubDate>Tue, 22 Sep 2026 17:27:48 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">NVIDIA</source>
      <description>As large language model (LLM) inference increasingly processes sensitive information and proprietary model context across personal, enterprise, and regulated...</description>
    </item>
    <item>
      <title>Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore</title>
      <link>https://aws.amazon.com/blogs/machine-learning/evaluate-skill-equipped-agents-with-strands-evals-and-amazon-bedrock-agentcore/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/evaluate-skill-equipped-agents-with-strands-evals-and-amazon-bedrock-agentcore/</guid>
      <pubDate>Tue, 22 Sep 2026 17:18:13 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>Skills let you encode domain-specific procedures as reusable, portable instructions for agents, but a fluent answer doesn&#x27;t prove the agent picked the right skill or followed it. Learn how to measure skill selection and instruction following with Strands Evals and Amazon Bedrock AgentCore Evaluations.</description>
    </item>
    <item>
      <title>Topology-Aware Workload Scheduling with NVIDIA Topograph</title>
      <link>https://developer.nvidia.com/blog/topology-aware-workload-scheduling-with-nvidia-topograph/</link>
      <guid isPermaLink="true">https://developer.nvidia.com/blog/topology-aware-workload-scheduling-with-nvidia-topograph/</guid>
      <pubDate>Tue, 22 Sep 2026 17:16:36 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">NVIDIA</source>
      <description>AI factories are power-limited systems that deliver maximum value when fully optimized. GPU workload placement is a key optimization. Poor workload placement...</description>
    </item>
    <item>
      <title>How Reactiv automates mobile commerce 80% faster with Amazon Bedrock AgentCore</title>
      <link>https://aws.amazon.com/blogs/machine-learning/how-reactiv-automates-mobile-commerce-80-faster-with-amazon-bedrock-agentcore/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/how-reactiv-automates-mobile-commerce-80-faster-with-amazon-bedrock-agentcore/</guid>
      <pubDate>Tue, 22 Sep 2026 15:46:07 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>Reactiv used Amazon Bedrock AgentCore to build a multi-agent AI Scheduler that autonomously refreshes Shopify merchants&#x27; mobile apps on a schedule, reducing merchant configuration time by 80% and getting to production 33% faster.</description>
    </item>
    <item>
      <title>Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI</title>
      <link>https://aws.amazon.com/blogs/machine-learning/right-size-generative-ai-endpoints-with-concurrency-sweeps-on-amazon-sagemaker-ai/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/right-size-generative-ai-endpoints-with-concurrency-sweeps-on-amazon-sagemaker-ai/</guid>
      <pubDate>Tue, 22 Sep 2026 15:35:53 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>Concurrency sweeps help you right-size a generative AI endpoint on Amazon SageMaker AI by systematically benchmarking it at increasing load levels. This post walks through deploying a model, running automated concurrency sweeps with the CreateAIBenchmarkJob API, and using the results to make data-driven capacity decisions about fleet size.</description>
    </item>
    <item>
      <title>How Trane gets building insights 60x faster with Amazon Bedrock AgentCore</title>
      <link>https://aws.amazon.com/blogs/machine-learning/how-trane-gets-building-insights-60x-faster-with-amazon-bedrock-agentcore/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/how-trane-gets-building-insights-60x-faster-with-amazon-bedrock-agentcore/</guid>
      <pubDate>Tue, 22 Sep 2026 15:30:34 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>In about four weeks, Trane Technologies built an AI-powered agentic solution on Amazon Bedrock AgentCore that reduced a 20-minute, multi-screen building diagnostic workflow to a 20-second natural language interaction, a 60x improvement in time-to-insight. This post shares the architectural approach and key design decisions behind the solution.</description>
    </item>
    <item>
      <title>How Tata Elxsi detects industrial safety risks in seconds on AWS</title>
      <link>https://aws.amazon.com/blogs/machine-learning/how-tata-elxsi-detects-industrial-safety-risks-in-seconds-on-aws/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/how-tata-elxsi-detects-industrial-safety-risks-in-seconds-on-aws/</guid>
      <pubDate>Tue, 22 Sep 2026 15:19:54 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>Learn how Tata Elxsi built IRIS, a real-time industrial safety platform on AWS. IRIS filters camera video at the edge, streams metadata through Amazon Kinesis, runs computer vision on Amazon SageMaker AI, and correlates detections into high-confidence alerts, detecting unsafe conditions in seconds instead of minutes.</description>
    </item>
    <item>
      <title>Extending public sector intelligence with Agentforce and AWS</title>
      <link>https://aws.amazon.com/blogs/machine-learning/extending-public-sector-intelligence-with-agentforce-and-aws/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/extending-public-sector-intelligence-with-agentforce-and-aws/</guid>
      <pubDate>Tue, 22 Sep 2026 15:17:45 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>Public sector agencies process large volumes of unstructured evidence, such as body camera footage and scanned documents. This post shows how to combine Amazon Bedrock Data Automation with the Model Context Protocol (MCP) to turn that data into structured insights and surface them through natural language queries in Salesforce Agentforce.</description>
    </item>
    <item>
      <title>Parallel cut research time and cost in half with GPT‑6 Astra</title>
      <link>https://openai.com/index/parallel-cuts-time-and-cost-with-astra</link>
      <guid isPermaLink="true">https://openai.com/index/parallel-cuts-time-and-cost-with-astra</guid>
      <pubDate>Tue, 22 Sep 2026 12:00:00 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">OpenAI</source>
      <description>GPT‑6 Astra allowed Parallel’s agents to research and synthesize labor-market data in half the time and at half the cost vs. prior models.</description>
    </item>
    <item>
      <title>Transformers now runs llama.cpp quants</title>
      <link>https://huggingface.co/blog/transformers-llama-cpp-quants</link>
      <guid isPermaLink="true">https://huggingface.co/blog/transformers-llama-cpp-quants</guid>
      <pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">Hugging Face</source>
      <description>Hugging Face published this.</description>
    </item>
    <item>
      <title>Priorities and principles for effective third party assessments</title>
      <link>https://openai.com/index/priorities-principles-third-party-assessments</link>
      <guid isPermaLink="true">https://openai.com/index/priorities-principles-third-party-assessments</guid>
      <pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">OpenAI</source>
      <description>OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.</description>
    </item>
    <item>
      <title>Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community</title>
      <link>https://huggingface.co/blog/omlx</link>
      <guid isPermaLink="true">https://huggingface.co/blog/omlx</guid>
      <pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">Hugging Face</source>
      <description>Hugging Face published this.</description>
    </item>
    <item>
      <title>How UK AISI and EvalEval Are Making Benchmark Results Reproducible</title>
      <link>https://huggingface.co/blog/evaleval-aisi</link>
      <guid isPermaLink="true">https://huggingface.co/blog/evaleval-aisi</guid>
      <pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">Hugging Face</source>
      <description>Hugging Face published this.</description>
    </item>
    <item>
      <title>Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton</title>
      <link>https://developer.nvidia.com/blog/simplifying-model-serving-across-multiple-gpus-with-nvidia-tensorrt-multi-device-integration-in-nvidia-dynamo-triton/</link>
      <guid isPermaLink="true">https://developer.nvidia.com/blog/simplifying-model-serving-across-multiple-gpus-with-nvidia-tensorrt-multi-device-integration-in-nvidia-dynamo-triton/</guid>
      <pubDate>Mon, 21 Sep 2026 21:51:14 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">NVIDIA</source>
      <description>The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability...</description>
    </item>
    <item>
      <title>How to Evaluate AI Agents From Tool Calls to Task Completion</title>
      <link>https://developer.nvidia.com/blog/how-to-evaluate-ai-agents-from-tool-calls-to-task-completion/</link>
      <guid isPermaLink="true">https://developer.nvidia.com/blog/how-to-evaluate-ai-agents-from-tool-calls-to-task-completion/</guid>
      <pubDate>Mon, 21 Sep 2026 21:05:37 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">NVIDIA</source>
      <description>When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and...</description>
    </item>
    <item>
      <title>Accelerating a ROS 2 Node with an AI Agent and NVIDIA Isaac ROS</title>
      <link>https://developer.nvidia.com/blog/accelerating-a-ros-2-node-with-an-ai-agent-and-nvidia-isaac-ros/</link>
      <guid isPermaLink="true">https://developer.nvidia.com/blog/accelerating-a-ros-2-node-with-an-ai-agent-and-nvidia-isaac-ros/</guid>
      <pubDate>Mon, 21 Sep 2026 19:07:33 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">NVIDIA</source>
      <description>GPU acceleration can speed up compute-intensive robotics workloads, but a fast CUDA kernel alone does not guarantee a fast ROS 2 graph. As messages move between...</description>
    </item>
    <item>
      <title>Benchmarking LLM Inference at Scale with AIPerf</title>
      <link>https://developer.nvidia.com/blog/benchmarking-llm-inference-at-scale-with-aiperf/</link>
      <guid isPermaLink="true">https://developer.nvidia.com/blog/benchmarking-llm-inference-at-scale-with-aiperf/</guid>
      <pubDate>Mon, 21 Sep 2026 18:45:07 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">NVIDIA</source>
      <description>You’re deploying a model on a system. It starts up, prompts are getting responses. Now the hard question: Is this fast? Your instincts might lead you to send...</description>
    </item>
    <item>
      <title>xAI’s Grok 4.6 is now available in Amazon Bedrock</title>
      <link>https://aws.amazon.com/blogs/machine-learning/xais-grok-4-6-is-now-available-in-amazon-bedrock/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/xais-grok-4-6-is-now-available-in-amazon-bedrock/</guid>
      <pubDate>Mon, 21 Sep 2026 18:30:34 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>xAI&#x27;s Grok 4.6 is now available in Amazon Bedrock: a frontier model for long-running agents, coding, and knowledge work, with a 500K token context window and four reasoning effort levels. It runs on both the bedrock-mantle and bedrock-runtime endpoints, with Converse API and cross-Region inference support.</description>
    </item>
    <item>
      <title>GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay</title>
      <link>https://arxiv.org/abs/2609.25001v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.25001v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:59:33 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.</description>
    </item>
    <item>
      <title>WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory</title>
      <link>https://arxiv.org/abs/2609.24984v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24984v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:57:06 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator&#x27;s limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.</description>
    </item>
    <item>
      <title>onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction</title>
      <link>https://arxiv.org/abs/2609.24983v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24983v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:56:53 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator locates the first inappropriate token and either picks a substitute from the model&#x27;s candidate tokens or types the correct text via free-form editing. The system then truncates everything after that position and continues generation from the corrected prefix, repeating this locate-correct-continue loop until a satisfactory response is obtained. This mechanism lets annotators precisely steer model outputs at low cost: a small controlled study suggests that onPanda reduces median annotation time by 52% over manual post-editing. Since the vast majority of tokens in the final response are generated by the model itself, the resulting data largely preserves the model&#x27;s sampling distribution and is well suited for constructing on-policy SFT and preference data. Furthermore, the token-level corrections recorded during annotation provide fine-grained supervision with precise positions and naturally paired positive--negative samples. onPanda also connects to external tools and harnesses, enabling interactive trajectory annotation in realistic environments. In addition, we release Panda-CVL, a dataset annotated with onPanda, together with a benchmark for token-level correction.</description>
    </item>
    <item>
      <title>LoRA-generating hypernetworks for efficient on-device LLM generative personalization</title>
      <link>https://arxiv.org/abs/2609.24979v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24979v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:55:48 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>On-device large language models (`LLMs&#x27;), e.g. running on mobile phones, are ripe for improvement via personalization. The limited compute resources of mobile devices impose limits on model scale and thus model quality, making any realizable quality gains highly impactful. At the same time, their personal nature (i.e., the close coupling to a particular user) means that a given on-device LLM tends to be used in similar, predictable patterns over the course of time. This paper presents a novel method for personalizing on-device LLMs. It trains a hypernetwork to map a user&#x27;s context tokens to a low-rank adaptation (`LoRA&#x27;) well-suited to that user. Once the trained common artifacts are deployed to users&#x27; devices, each user uses the hypernetwork to synthesize (entirely on device) a personalized LoRA. This approach blends the benefits while avoiding the drawbacks of two existing approaches to LLM customization: in-context learning (`ICL&#x27;) and parameter-efficient fine-tuning (`PEFT&#x27;). Like ICL (and unlike PEFT), the on-device phase of our approach is computationally feasible, requiring only forward passes through neural networks. Like PEFT (and unlike ICL), our approach modifies the `target&#x27; base LLM via weights (the LoRA), avoiding negative consequences (e.g. increased latency) associated with extending the input sequence. Our approach is particularly well-suited to the mobile device regime. Apart from the on-device compute and latency benefits mentioned, it also requires minimal additional storage, as internally its architecture partly leverages the same LLM weights as belong to the target LLM to be personalized. We demonstrate the benefits of LoRA-generating hypernetworks on several representative personalization datasets, comparing against baselines like ICL and PEFT. Of note, our personalization experiments focus on more challenging and less studied long-form text generation tasks.</description>
    </item>
    <item>
      <title>DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation</title>
      <link>https://arxiv.org/abs/2609.24976v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24976v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:55:24 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.</description>
    </item>
    <item>
      <title>Harness-Zero: Harness Distillation via Agent-as-Harness</title>
      <link>https://arxiv.org/abs/2609.24974v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24974v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:55:20 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time guidance and transferring the behaviors it induces into model weights, so that its gains survive under a single fixed target harness. The challenge is that the two harnesses differ in action space and available information, so guidance from the optimized harness cannot serve directly as supervision for the target one. We introduce Harness-Zero, which enables harness distillation through agent-as-harness. Guided by the optimized harness, a harnessing agent corrects student responses before execution in the target harness&#x27;s action space, turning harness guidance into training demonstrations. Fine-tuning on the resulting trajectories internalizes harness-induced behavior into the model, so the specialized harness can be removed at deployment. Our experiments spanning knowledge work, tool use, and science domains show that: (1) For frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness. (2) With the specialized harness removed at deployment, Harness-Zero improves the base model&#x27;s macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached. (3) Harness-Zero recovers harness-induced behaviors absent from the base model, with 82.3% average recovery across 28 patterns in the three domains.</description>
    </item>
    <item>
      <title>RRSI: Regularized Recursive Self-Improvement of Agent Harnesses</title>
      <link>https://arxiv.org/abs/2609.24972v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24972v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:54:49 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>An LLM agent&#x27;s capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.</description>
    </item>
    <item>
      <title>DolphinBench: Mapping the Pareto Frontier of Agent Memory</title>
      <link>https://arxiv.org/abs/2609.24971v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24971v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:54:35 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent&#x27;s task completion. DolphinBench includes three knowledge-work personas with roughly 500k tokens of user messages per persona and evaluates agents on tasks that depend on information from that history. We verify all 200 tasks per persona by running an agent with and without the relevant history, requiring success with it and failure without it. Finally, we require all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically. No existing memory benchmark combines all three. The dataset and evaluation code are available at https://dolphinbench.ai.</description>
    </item>
    <item>
      <title>Rare Event Estimation via Iterative Unalignment</title>
      <link>https://arxiv.org/abs/2609.24969v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24969v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:53:38 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We study the problem of estimating the probability of rare events that arise from stochastic variation in the agent&#x27;s own actions. Estimating this type of risk requires searching over the combinatorially vast space of trajectories. Naive Monte Carlo is computationally prohibitive in this regime, and constructing effective importance sampling (IS) proposals requires coordinated changes to a context-dependent chain of conditional distributions. We develop a new IS method that perturbs the original model&#x27;s weights to construct the proposal. The proposal is itself a differentiably parameterized language model, enabling gradient-based search over weight space. We formulate an objective that combines a differentiable surrogate for event amplification and an adaptive regularization scheme that dynamically balances amplification against estimator stability. We evaluate our approach on $\sim$120M and $\sim$2.6B models across three event families spanning 300+ rare events as rare as $10^{-9}$, with reference probabilities computed with $&lt;10\%$ relative standard error. In our most verifiable settings, we observe that our IS estimator achieves over $800\times$ compute-weighted efficiency gains over naive Monte Carlo for events with probabilities lower than $10^{-7}$. Our implementation is available at https://github.com/namkoong-lab/iterative-unalignment.</description>
    </item>
    <item>
      <title>Emergent Collusion in Long-Horizon LLM Agent Interaction</title>
      <link>https://arxiv.org/abs/2609.24967v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24967v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:52:48 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other&#x27;s work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of reward structure, the verification feedback agents receive, and their interaction history. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.</description>
    </item>
    <item>
      <title>Learning Physics from an Imperfect Ancestor</title>
      <link>https://arxiv.org/abs/2609.24947v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24947v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:42:17 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Neural operators evaluate parametric partial differential equations cheaply but degrade sharply outside their training distribution. Physics-informed neural networks avoid dependence on labeled data, yet their optimization can be basin-fragile: when the governing residual admits multiple solutions, a PINN trained from scratch may converge to a physically incorrect state despite achieving a small residual. We show that these failure modes can be addressed jointly: an imperfect NO provides the structural prior needed to place a PINN in the correct solution basin, while the PDE residual refines the solution beyond the operator&#x27;s accuracy. We introduce a three-stage framework that freezes the spatial basis of a physics-informed NO, extrapolates its solution branch to an out-of-distribution parameter using a polynomial continuation prior, and distills the resulting field into a fresh PINN. The NO need not be accurate at the target; it transfers solution-branch information, while PDE residual minimization in the PINN governs convergence. We evaluate the framework on three nonlinear PDEs: 1D viscous Burgers, 2D steady Allen-Cahn near a pitchfork bifurcation, and 2D steady lid-driven cavity flow. For Allen-Cahn, where the trivial solution satisfies the PDE residual exactly, a standard PINN collapses to the trivial zero branch, whereas distillation from the crude extrapolated operator recovers the non-trivial branch that matches the finite-difference reference. For the lid-driven cavity, extrapolating to a Reynolds number of Re = 3200 accelerates convergence to the correct physical state, achieving competitive accuracy using fewer parameters and optimization steps than recent literature baselines. These results establish a simple principle: an NO need not accurately predict the solution to be useful; it only needs to identify the correct basin from which PINN optimization can recover it.</description>
    </item>
    <item>
      <title>Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization</title>
      <link>https://arxiv.org/abs/2609.24942v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24942v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:39:51 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>A model generalizes outside its training distribution only when it computes a representation structurally equivalent to the generating mechanism, not an approximation fitted to it. Such equivalence is necessary for exactness in and out of distribution, and extrapolation is governed by this exactness at inference, whatever its realization. Tensor Logic shows this: a zero-temperature contraction is equivalent to discrete logic, deducing in place with no artefact extracted, its tensors Boolean, its embeddings orthonormal, only its arithmetic continuous. Lacking infinite recursion it reaches Datalog, not Prolog, and though exact over closed domains it needs external memory to bind a novel entity. The criterion needs neither a discrete representation nor an extracted expression, and constrains inference, not training: an exact marginal in $[0,1]$ passes, a Neural Network thresholded to a hard label does not. Logic Tensor Networks fail it, while differentiable ILP and Tensor Logic at $T=0$ pass. Piecewise-affine extrapolation divergence and an inability to bind novel entities are two faces of a shortfall in exact representability. For hybrid architectures, a propagation rule follows: the output inherits the bounds of every fitted estimator on its path, explaining which axes fail in equivariant models and the ARC-AGI induction/transduction split. Only an exact hypothesis class certifies what the training data leave underdetermined: on a law-derived partition it finds the $56.3\%$ of distant queries that are answerable, which ensembles meet with false confidence and distance metrics rank backwards. Common inductive biases, from symmetries to memory, reach exactness only because humans inject them, an argument for inducing exact representations rather than fitting surrogates whose residuals, even at the arithmetic floor in training, diverge outside the data and compound under composition.</description>
    </item>
    <item>
      <title>Linguistic Features for Interpretable Textual Entailment</title>
      <link>https://arxiv.org/abs/2609.24932v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24932v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:26:30 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Despite the success of neural models in natural language processing, their black-box nature limits interpretability and conceals the linguistic phenomena underlying their predictions. We present SLITE, an explainable hybrid model for Recognizing Textual Entailment that integrates two complementary layers of semantic analysis: a structural-relational layer, based on semantic compatibility and incompatibility between compositional entities, and a distributional-informational layer, based on structured patterns of information change between embedding-based representations of the premise and the hypothesis. We propose 17 features that combine entity-level semantic relations, polarity-sensitive lexical matching, and alignment measures over semantic sub-representations of the similarity matrix, including measures based on entropy and transfer entropy. A logistic regression trained on these features achieves an accuracy of 83% on three-class SICK and 96% on SICK-CE, outperforming IsoLex by 4 percentage points and falling within 2 percentage points of RoBERTa with a fraction of its computational complexity. Ablation studies and SHAP analysis confirm that structural-relational features are the primary drivers of classification, while distributional-informational features provide essential complementary contributions, particularly for detecting neutrality and contradiction. Our results demonstrate that further exploration of hybrid approaches is a viable and scientifically productive alternative to massive neural architectures, and we hope they will strengthen the dialogue between linguistic theory and computational modeling of inference</description>
    </item>
    <item>
      <title>Conformalized Quantile Regression and Minimax Limits of Fixed-Score Calibration under Known Covariate Shift</title>
      <link>https://arxiv.org/abs/2609.24929v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24929v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:24:00 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>In this paper, we study nonasymptotic $L^p$ error bounds for interval length and conditional coverage in split conformalized quantile regression (CQR). Our bounds rely on local regularity conditions and accuracy guarantees for the estimated quantiles. We further instantiate our bounds for quantile regression with sparse ReLU neural networks. We also consider covariate shift, where the calibration and test covariates have different distributions, and derive nonasymptotic bounds for this setting. We obtain matching minimax upper and lower bounds in expectation for two constructed fixed-score calibration benchmarks under known covariate shift. The bounds match for every $p\in[1,\infty]$ in the scalar problem and for finite $p$ in the $K$-threshold problem; for the latter, a high-probability minimax lower bound holds for every $p\in[1,\infty]$.</description>
    </item>
    <item>
      <title>Et Tu, Brute? Economic Misalignment in Personal AI Agents</title>
      <link>https://arxiv.org/abs/2609.24927v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24927v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:22:44 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Personal AI agents make recommendations and take actions on people&#x27;s behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user&#x27;s personal context, e.g., their email inbox and a structured profile of personal attributes, with the intention of making an optimal, personalized decision for the user. We show that by simply providing this personal context, the agent steers recommendations based on inferred wealth, without being explicitly instructed to do so. In a suite of 325K experiments on 13 agents across three types of economic decisions (flights, health insurance, and graduate programs), we find that 8 models systematically choose more expensive options for wealthier users when requests are identical. This steering continues even when it directly goes against the user&#x27;s stated objective: when explicitly instructed to find the cheapest option, some agents still act on the wealth profile they have inferred. It also occurs when wealth is inferred from ambient data, such as emails unrelated to the task. And it persists under privacy controls that block specific attributes: blocking financial attributes largely removes the disparity, but blocking other attributes leaves it unchanged and can increase it by up to 40% for insurance, as agents rely on the remaining signals to infer wealth. Larger and more capable models are no better; Claude Opus 4.8 shows the largest effect. We term this misalignment &quot;adversarial delegation&quot;, in which the very conditions that make a personal AI agent useful - access to personal information - enable it to act against the user&#x27;s interests.</description>
    </item>
    <item>
      <title>BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction</title>
      <link>https://arxiv.org/abs/2609.24921v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24921v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:18:19 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Scientific weak signals are early, low-visibility research directions that later become central to mature scientific topics, yet existing resources such as trend tracking, citation forecasting, and foresight reports rarely provide validated reference sets that link concrete early precursors to later paradigms. We introduce BackTrend, a retrospective benchmark in which, given a mature target topic and a temporal evidence constraint, systems must recover two types of precursors: problem-space signals, underrecognized research problems, and solution-space signals, emerging methods for known problems. BackTrend contains 25 mature target topics in artificial intelligence and machine learning and 66 human-validated weak signals, reconstructed from large-scale literature by grounding each candidate in its 2019-2024 publication-frequency trajectory. We evaluate frontier LLMs, RAG systems, and agentic research systems using semantic matching and coverage-based metrics. Current systems often generate plausible but misaligned precursors, exhibiting topic drift, granularity mismatch, near-miss matching, and incomplete coverage; the strongest system achieves only 10.1% F1, while Coverage10 reaches at most 18.5% of the reference signals. Our budget analyses show that additional retrieval and web-search evidence can improve performance up to a moderate budget, but does not by itself close the substantial performance gap.</description>
    </item>
    <item>
      <title>SocioVerse2: A Longitudinal Dynamic Social Simulation Framework under a Human-AI Co-evolutionary Paradigm</title>
      <link>https://arxiv.org/abs/2609.24911v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24911v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:14:09 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Social simulation offers the social sciences an experimental instrument that the real world cannot supply, and generative agents have transformed it by acting as silicon samples that unite agent-based modeling with real behavioral data. Existing platforms verify collective behavior, align simulated populations with real societies in cross-sections, and employ autonomous agents for the research process. However, two social science requirements remain without systematic support: intervention in the content of a simulation and the researcher&#x27;s control over the process that produces it. We present SocioVerse2, which extends SocioVerse 1.0 into a human-AI co-evolutionary paradigm built from two loops and one infrastructure. The longitudinal simulation loop simulates the target population with evolving environments and forks counterfactual branches via interventions. The controllable research loop takes the study itself as an editable state and updates state versions via controllable editing. The social science agentic infrastructure carries both loops through composable skills with researcher checkpoints, a population service over five persona pools, and an environment service over 21 real-world signal sources with point-in-time guarantees. We validate SocioVerse2 across three case families and seven case studies, from reproducing canonical agent-based models to modeling policy processes on real records and nowcasting macro-economic indices beyond the response model&#x27;s knowledge cutoff. With the human-AI co-evolutionary paradigm, these cases go beyond system demonstrations to become substantive studies that investigate frontier questions in their respective disciplines. Code, data services, and a workbench are released as open-source resources.</description>
    </item>
    <item>
      <title>Visuomotor Robotic Pruning in Planar Orchards Using Hybrid Reinforcement Learning</title>
      <link>https://arxiv.org/abs/2609.24906v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24906v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:09:58 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Dormant tree pruning is labor-intensive yet essential for maintaining modern high-productivity fruit orchards. In this work, we focus on pruning of modern planar tree training systems - V-Trellis apples and UFO cherries - where trunks and primary branches are trained into approximately planar walls. We introduce an end-to-end pipeline to learn a closed-loop visuomotor controller for robotic pruning. This controller is trained entirely using simulation and synthetically generated data and deployed in real orchards in a zero-shot manner. The pipeline comprises synthetic generation of planar orchard tree meshes, construction of a physics-based orchard simulator, automated collection of successful pruning trajectories via motion planning, and policy learning with a novel hybrid reinforcement-learning algorithm that combines offline demonstrations with online simulated rollouts. The controller uses optical-flow inputs from a wrist-mounted camera - avoiding the need for full 3D-reconstruction - and continuously guides the cutter through cluttered branch environments to a specified cutpoint with correct tool orientation. In exhaustive simulated task-space evaluations over 3,000 pruning points, the policy attains 49.9% success on V-Trellis apples and 46.0% on UFO cherries. We validate the learned controller across 38 physical trials - comprising 28 outdoor field trials in commercial and experimental orchards and 10 indoor laboratory tests - demonstrating zero-shot sim-to-real transfer. The learned policy also outperforms a classical RRT-Connect baseline on physical hardware in laboratory trials.</description>
    </item>
    <item>
      <title>ToneCL: Contrastive Learning for Few-Shot Syllable-Level Tone Classification</title>
      <link>https://arxiv.org/abs/2609.24903v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24903v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:09:21 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Tone languages constitute over 50-70% of the world&#x27;s languages, but the vast majority are low-resource, lacking the large transcribed corpora needed for automatic tone classification. Existing datasets are typically collected at the sentence level, whereas field linguists require fine-grained syllable-level annotations. We propose ToneCL, a lightweight contrastive learning framework for few-shot syllable-level tone classification. We simulate low-resource conditions on Mandarin and Vietnamese, limiting labeled data to tens of examples per tone class. ToneCL is pretrained on unlabeled speech with augmentations that preserve tonal identity, then fine-tuned on few-shot examples. Experiments show our method consistently outperforms baselines, achieving 91.6% on six-speaker Mandarin at 10 shots. Cross-lingual transfer is also effective: pretraining on Vietnamese and fine-tuning on Mandarin reaches 91.0\% accuracy at 10 shots. Ablation confirms that frequency band rejection is the most critical augmentation.</description>
    </item>
    <item>
      <title>Human-LLM Deliberation as Interactive Proof: Conditions for Verifiability Without Transparency</title>
      <link>https://arxiv.org/abs/2609.24895v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24895v1</guid>
      <pubDate>Mon, 21 Sep 2026 17:00:15 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>When an LLM supplies an argument that a user could not readily construct, how can the user decide whether to accept its claim? Inspired by interactive proofs, we model human-LLM deliberation as an interaction between a prover with unrestricted internal search and a resource-bounded human verifier. The verifier requests and checks supporting details without access to the LLM&#x27;s internal state. Passed checks accumulate evidence toward an acceptance threshold. We prove anytime-valid soundness against adaptive provers: the probability of ever accepting a false claim is at most a chosen error level, provided the task supplies bounds on false passes and human checking errors that remain valid after every relevant history. A finite-horizon completeness bound additionally requires bounds on the adequacy of honest responses and sufficient diagnostic progress. Further checks can strengthen the evidence for acceptance, but each requires another adequate response and reliable human effort. Whether this tradeoff permits certification depends on the verifier&#x27;s effort budget, cognitive load, expertise, and fatigue. We identify conditions under which the supplied bounds certify a specified sequence of local checks but not a specified global check under the same resource budgets.</description>
    </item>
    <item>
      <title>SLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models</title>
      <link>https://arxiv.org/abs/2609.24894v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24894v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:59:06 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv</source>
      <description>Whole-slide pathology images (WSIs) contain gigapixel-scale visual content, creating a major scalability challenge for slide-level multimodal large language models (MLLMs). Existing approaches process thousands of patch tokens and typically apply compression only after slide encoding, leaving multimodal attention computationally expensive. We introduce SLICEChat, a slide-level MLLM that integrates progressive token pruning within a hybrid Mamba--Transformer slide encoder. Mamba layers enable efficient long-range propagation, while Transformer layers preserve global interactions as the sequence is progressively shortened. Between stages, language-supervised, region-aware pruning removes spatially coherent low-utility regions under a controlled keep-rate schedule, producing compact slide representations before multimodal fusion. On SlideBench VQA, SLICEChat achieves 79.84% accuracy on TCGA and 59.09% on BCNB cohorts, outperforming prior slide-level pathology MLLMs, and achieves the highest overall WSI-Bench metrics. It also provides competitive memory usage and the inference latency among the evaluated models. These results demonstrate accurate and computationally efficient multimodal reasoning over gigapixel WSIs.</description>
    </item>
    <item>
      <title>OSWorld-Pro: Process-based Evaluation for Computer Use Agents</title>
      <link>https://arxiv.org/abs/2609.24890v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24890v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:55:24 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.CL</source>
      <description>Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agents fail in various tasks, obfuscating critical insight for subsequent improvement. For instance, agents that err during keyboard inputs would require a different mitigation strategy from those that fail to precisely provide click-based inputs on the graphical UI. We introduce OSWorld-Pro: a set of over 300 tasks containing over 2800 subgoals to enable the procedural evaluation of CUAs grounded in over 67,000 human annotations. We use robust human-aligned LLM-Judges to evaluate the fulfillment of OSWorld-Pro subgoals and thereby reveal the progress that models make throughout a series of sequentially dependent subgoals. Our findings reveal that OSWorld-Pro is challenging even for state-of-the-art LLMs, with top performers like Claude Opus 5 achieving only 75.7% vs. 83.4% on OSWorld. Furthermore, we identify critical process-focused failure modes of various models (e.g. subgoal-irrelevant actions and click-based mistakes) to provide insights to improve performance and efficiency of CUAs.</description>
    </item>
    <item>
      <title>The Copy Ceiling: An Input-Exposure Control for Ontology-Grounded Generation over Curated Corpora</title>
      <link>https://arxiv.org/abs/2609.24885v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24885v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:53:00 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.CL</source>
      <description>When a language model answers from a curated corpus via graph-based retrieval, a large grounding uplift does not establish reasoning over the retrieved structure: the context may already expose the gold answers. We propose exposure accounting, which classifies each gold item by whether the shown context exposes it and whether the answer recovers it. Its scalar reference is the copy ceiling, the recall a verbatim copy of the context achieves; signed gain over copy measures the model&#x27;s recall relative to this deterministic, judge-free baseline. Across ten models, unaided recall averages 0.26 and grounded recall 0.92, yet gain over copy is uniformly negative (-0.067 to -0.022). Of 11,360 gold-item observations, representing 1,136 target instances evaluated under ten models, only three unexposed items receive lexical credit. A stratified model-judged audit of 423 observations, with a symmetric quotation-verification policy, estimates that 97.1% of credited items assert the requested relation; all three unexposed credits fail relational adjudication. On targets the scaffold does not expose, lexical recovery falls from 0.121 unaided to 0.004 grounded; adjudication validates 71 of the 92 unaided credits and none of the three grounded credits, without establishing full-frame relational recovery rates. Rephrasing questions outside the graph&#x27;s title vocabulary reduces exposure from 0.964 to 0.328, while an absence-triggered fallback activates on only 2 of 506 questions. A paired production study improves judged quality by +0.27 pooled, but negative controls do not establish content specificity beyond a well-formed on-corpus block. These results support exposure accounting as a standing control for corpus-derived evaluations. The accounting distinguishes exposed-item omissions from beyond-exposure recoveries; it does not determine whether reasoning occurred.</description>
    </item>
    <item>
      <title>Learning Prognostic Variables for AI Convective Parameterizations via Symbolic Distillation</title>
      <link>https://arxiv.org/abs/2609.24882v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24882v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:50:03 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.LG</source>
      <description>Hybrid AI-physics climate modeling aims to improve coarse (~100km-resolution) Earth system models by learning to parameterize subgrid processes from high-fidelity data. However, this so far mostly involves local-in-time, diagnostic parameterizations, in which the subgrid state depends only on the current coarse state with no memory of previous states, which is unrealistic for processes such as convection that have intrinsic persistence. To address this, we enhance local-in-time parameterizations by learning prognostic variables that compactly carry important, additional past information where no explicit sub-grid information is available. First we compress past information into a low-dimensional latent space using an autoencoder, which then informs a neural network trained to parameterize targeted subgrid-scale processes. We then replace the autoencoder with symbolic equations that govern the time evolution of the latent variables, yielding additional prognostic memory variables that can be integrated alongside the resolved atmospheric state. We evaluate this approach on two systems: the Lorenz-96 model (online) and surface precipitation from high-resolution atmospheric simulations (offline). A forced multivariate linear ordinary differential equation recovers most of the added value achieved by the autoencoder-based approach in both experiments. Benchmarked against diagnostic parameterizations without memory, our memory-informed approach improves climate statistics and temporal structure, including a realistic diurnal cycle of tropical land precipitation.</description>
    </item>
    <item>
      <title>Run Positron on Amazon SageMaker AI for data science workflows</title>
      <link>https://aws.amazon.com/blogs/machine-learning/run-positron-on-amazon-sagemaker-ai-for-data-science-workflows/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/run-positron-on-amazon-sagemaker-ai-for-data-science-workflows/</guid>
      <pubDate>Mon, 21 Sep 2026 16:34:21 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>Positron, Posit&#x27;s IDE for data science, now runs on Amazon SageMaker AI. This post shows how a data scientist explores an Amazon Athena table, validates features in R, trains an XGBoost model in Python, deploys a real-time SageMaker AI endpoint, and reports results with Quarto, all in one governed SageMaker Studio Space.</description>
    </item>
    <item>
      <title>When Tomorrow Becomes Today: Self-Evolving Policies for Agentic Time-Series Forecasting</title>
      <link>https://arxiv.org/abs/2609.24862v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24862v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:33:59 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.LG</source>
      <description>Agentic time series forecasting concerns systems whose underlying mechanisms evolve, making the relative effectiveness of numerical models, reasoning strategies, and intervention rules inherently time-varying. Consequently, a time series agent must adapt the forecasts it produces and the orchestration policy that determines which components to trust and how to coordinate them. The deployment process naturally provides supervision for this adaptation as forecast horizons elapse and realized targets reveal the effectiveness of earlier decisions. Committing all numerical expert forecasts and candidate agent paths before target observation allows each realized outcome to evaluate the entire alternative set, providing delayed feedback without additional annotation. However, existing time series agents primarily incorporate prior experience through forecast refinement, reflection, or retrieval, without systematically converting realized outcomes into persistent updates to the joint orchestration policy governing later origins. To exploit this delayed feedback systematically, we introduce TimEvolve, a frozen-backbone time series agent that converts each realized outcome into persistent joint updates of expert trust, agent path selection, and intervention strength. A temporally ordered predict, reveal, and update protocol applies this feedback to subsequent forecasts. Experiments across eight Time-MMD domains show that TimEvolve achieves the best average MSE and MAE ranks among fifteen methods and the lowest errors on both metrics in seven domains. These results demonstrate the value of learning forecasting policies from the futures encountered during deployment.</description>
    </item>
    <item>
      <title>Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection</title>
      <link>https://arxiv.org/abs/2609.24855v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24855v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:30:38 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.CL</source>
      <description>Argumentative component detection (ACD) is a core subtask of Argument(ation) Mining (AM) and one of its most challenging aspects, as it requires jointly delimiting argumentative spans and classifying them into components such as claims and premises. While research on this subtask remains relatively limited compared to other AM tasks, most existing approaches formulate it as a simplified sequence labeling problem, component classification, or a pipeline of component segmentation followed by classification. In this paper, we propose ITFACD, a novel approach based on instruction-tuned Large Language Models (LLMs) using compact instruction-based prompts, and reframe ACD as a language generation task, enabling arguments to be identified directly from plain text without relying on pre-segmented components. Experiments on standard benchmarks show that our approach achieves higher performance compared to state-of-the-art systems. To the best of our knowledge, this is one of the first attempts to fully model ACD as a generative task, highlighting the potential of instruction tuning for complex AM problems. Our code and the datasets used are openly available in the following GitHub repository.</description>
    </item>
    <item>
      <title>How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore</title>
      <link>https://aws.amazon.com/blogs/machine-learning/how-benchling-secured-multi-tenant-ai-agents-with-amazon-bedrock-agentcore/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/how-benchling-secured-multi-tenant-ai-agents-with-amazon-bedrock-agentcore/</guid>
      <pubDate>Mon, 21 Sep 2026 16:27:34 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>Learn how Benchling built a defense-in-depth security architecture to run untrusted, AI agent-generated scientific code across thousands of life sciences tenants using Amazon Bedrock AgentCore Code Interpreter in VPC mode, combined with Amazon Route 53 Resolver DNS Firewall and VPC endpoint policies to block data exfiltration, including through DNS.</description>
    </item>
    <item>
      <title>Reducing medical claims review time with AI on AWS: The EXL Medical IDP solution</title>
      <link>https://aws.amazon.com/blogs/machine-learning/reducing-medical-claims-review-time-with-ai-on-aws-the-exl-medical-idp-solution/</link>
      <guid isPermaLink="true">https://aws.amazon.com/blogs/machine-learning/reducing-medical-claims-review-time-with-ai-on-aws-the-exl-medical-idp-solution/</guid>
      <pubDate>Mon, 21 Sep 2026 16:24:40 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">AWS Machine Learning</source>
      <description>EXL built an AI-powered Medical intelligent document processing (IDP) solution on AWS, combining IDP with domain-specific large language models on Amazon SageMaker and Amazon Bedrock to extract, summarize, and query medical records at enterprise scale and cut claims review time from over 100 minutes per case.</description>
    </item>
    <item>
      <title>PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control</title>
      <link>https://arxiv.org/abs/2609.24840v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24840v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:20:25 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.LG</source>
      <description>Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker&#x27;s capabilities, leaving recovery and physical execution largely to the tracker. Action-only diffusion generates actions directly but lacks an explicit future-state trajectory for test-time motion objectives. Joint state-action diffusion provides this representation, yet representative controllers often depend on privileged full-body states, and support for learned behavior selection and test-time motion steering remains fragmented. We present PredActor, a predictive action diffusion policy that brings these complementary steering capabilities into one directly executed policy using proprioceptive observations. Conditioned on proprioceptive history and optional task context, PredActor jointly generates executable actions and an internal future-state trajectory. Classifier-free guidance strengthens text-conditioned behavior, while classifier guidance steers predicted states toward test-time objectives. Only actions are executed, without a separate motion-reference tracker or externally estimated full-body states as policy inputs. In simulation, PredActor reaches all 15 destination targets and achieves a text retrieval score of 0.580, compared with 0.373 for conditional action diffusion, with similar observed disturbance survival. To make this guided policy practical onboard, rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, both below the 20 ms control period. We deploy PredActor on a Unitree G1; evaluations across simulation and physical hardware demonstrate text-conditioned motion, disturbance response, joystick control, and semantic interpolation.</description>
    </item>
    <item>
      <title>G-NAC: Graph Neural Automata Clustering via Emergent Domain Formation</title>
      <link>https://arxiv.org/abs/2609.24823v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24823v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:11:47 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.LG</source>
      <description>We introduce Graph Neural Automata Clustering (G-NAC), an unsupervised clustering method in which observations interact as cells on a fixed neighborhood graph. A shared recurrent graph-neural cellular rule evolves latent domain states through local interactions, which are converted into a rank-based spectral affinity for partitioning. Across 73 clustering tasks from 57 benchmark datasets, G-NAC achieved a mean adjusted Rand index (ARI) of 0.7951, comparable to Genie at 0.7941 and higher than the other evaluated baselines. Empirical training time and GPU memory scaled approximately linearly from 5,000 to 100,000 nodes. Learned transition rules also transferred from smaller source graphs to independent 100,000-node samples generated under matched conditions. These results demonstrate a recurrent graph-clustering formulation while identifying dependencies on graph quality, readout design, and source-target similarity.</description>
    </item>
    <item>
      <title>The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts</title>
      <link>https://arxiv.org/abs/2609.24821v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24821v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:11:01 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.CL</source>
      <description>The Linear Representation Hypothesis associates high-level concepts with directions in language models, but it remains unclear how these concept-related linear structures are organized within the model. We propose the Answer-Basin Representation Hypothesis: the probability measure induced over answers by the model&#x27;s continuation distribution organizes these linear structures, with its statistics represented along linear directions shared across questions. All continuations yielding the same answer form an answer basin, whose mass is their total probability. These basin masses define the pushforward probability measure over answers. We posit that concept-related linear structure emerges from differences in the answer measure rather than being determined by changes in concept labels. Experiments across models and tasks link concept-consistent effects and their reversals in probing and steering to the alignment between concept labels and the answer measure.</description>
    </item>
    <item>
      <title>MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents</title>
      <link>https://arxiv.org/abs/2609.24812v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24812v1</guid>
      <pubDate>Mon, 21 Sep 2026 16:04:26 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.CL</source>
      <description>Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, and collaborative work, are inherently multi-speaker. Supporting these settings introduces challenges that are largely absent from one-on-one interaction. We introduce the Multi-Speaker Interaction Benchmark (MSI-Bench) for evaluating multi-speaker voice interaction. Each test case is a short multi-party multi-turn audio scene with participant context, expected tool calls, and atomic rubrics. The benchmark targets three capability families: multi-speaker memory, multi-speaker instruction following, and multi-speaker reasoning. It comprises 1,152 test cases, evenly split between Mandarin Chinese and English (576 each). The strongest configuration on each split passes all rubrics on only 66.8% of English and 54.5% of Mandarin cases, and the strongest open-weight configuration on 34.0% and 19.3%. Failure analysis separates perception from reasoning: open-weight models are bottlenecked by the multi-speaker audio front-end, while frontier systems still fail speaker-scoped decision making on clean transcripts---and models across the board often respond when no one has addressed them. These results identify speaker-grounded perception, speaker-scoped decision making, and conversational restraint as concrete targets for future voice agents.</description>
    </item>
    <item>
      <title>When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs</title>
      <link>https://arxiv.org/abs/2609.24799v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24799v1</guid>
      <pubDate>Mon, 21 Sep 2026 15:59:06 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.CL</source>
      <description>Post-training quantization (PTQ) enables efficient deployment of large language models, and PTQ methods are usually optimized and evaluated with generic reconstruction, perplexity, or answer accuracy. But in explanation-critical domains, preserving only the final answer may be insufficient, since users may also inspect generated rationales to judge whether a prediction is trustworthy. We study this issue in medical multiple-choice question answering, where rationales should provide evidence that supports the selected answer. We propose an explanation-aware objective for transformation-based PTQ. Our method builds an offline faithfulness cache from full-precision teacher rationales and uses it during optimization to preserve answer-supporting evidence tokens and evidence-conditioned answer behavior. We instantiate it on OSTQuant under W4A4KV4 quantization and evaluate four 7B--8B medical and instruction-tuned LLMs on MedExQA, MedExpQA, and ChallengeClinicalQA. While a same-calibration OSTQuant baseline preserves task accuracy, it can substantially weaken answer-supporting rationales. Our objective is to preserve the full-precision model&#x27;s answer-supporting behavior rather than improve gold-label accuracy, and our method better preserves the full-precision model&#x27;s answer behavior and rationale-to-answer support. These results suggest that PTQ for explanation-critical settings should evaluate preservation of answer-supporting evidence, not only answer accuracy. Code and evaluation scripts are available at https://github.com/dut0817/EAQuant.</description>
    </item>
    <item>
      <title>Detecting Agitation Before Behavioral Escalation in Autistic Youth Through Multimodal Wearable Sensing</title>
      <link>https://arxiv.org/abs/2609.24791v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24791v1</guid>
      <pubDate>Mon, 21 Sep 2026 15:50:40 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.LG</source>
      <description>Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, vocalization, and autonomic arousal. Its signs are subtle and individualized, and its autonomic components are invisible without instrumentation. We collected upper-body movement from inertial measurement units, physiology from a wrist-worn device, and vocalizations from lapel microphones across 30 clinician-led sessions with 15 autistic youth, paired with expert behavioral annotations. We adapt four pretrained foundation models, one per modality, project each to a shared 128-dimensional space, and fuse them into a single group model. The model detected agitation with an area under the ROC curve of 0.724 at the clinician-annotated onset (within-participant permutation p=0.0005), declining to 0.608 at 30,s before onset. Thirteen of fifteen participants were above chance. A from-scratch configuration reached only 0.58, while frozen and fine-tuned features performed comparably (0.71 and 0.72). Audio contributed most of the signal, and a watch-only configuration stayed near chance. Individualized agitation is therefore detectable, including in unannotated windows preceding the annotated onset, using foundation-model transfer with one shared model rather than one per child.</description>
    </item>
    <item>
      <title>XSQ-AST: An Explainable Audio Spectrogram Transformer Framework for Localising Synthetic Speech Artifacts</title>
      <link>https://arxiv.org/abs/2609.24770v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24770v1</guid>
      <pubDate>Mon, 21 Sep 2026 15:38:11 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.LG</source>
      <description>Localising artifacts in synthetic speech remains challenging, as most evaluation methods yield only global quality scores. This paper presents XSQ-AST, a framework that combines the SQ-AST speech quality model with WhisperX phoneme alignment and multiple saliency methods to produce temporally localised artifact diagnostics without model retraining. Saliency maps are projected onto continuous distributions via kernel density estimation and onto phoneme boundaries via phoneme-discretised saliency maps. A 40-participant listening test validated the framework across five perceptual dimensions. Attention Rollout, Attention Flow and an adapted GradCAM produced temporal distributions that correlated with listener highlights, with different methods best suited to different artifact types. An AUC-ROC analysis confirmed discrimination above chance.</description>
    </item>
    <item>
      <title>Inference of Unknown Dynamical Components Using Next Generation Reservoir Computing: From Chaotic Systems to Climate Data</title>
      <link>https://arxiv.org/abs/2609.24754v1</link>
      <guid isPermaLink="true">https://arxiv.org/abs/2609.24754v1</guid>
      <pubDate>Mon, 21 Sep 2026 15:21:41 +0000</pubDate>
      <source url="https://alienlifeai.com/companies.xml">arXiv cs.LG</source>
      <description>We investigate next generation reservoir computing (NGRC) as a data-driven approach for inferring unseen components of dynamical systems. We compare NGRC with traditional reservoir computing (RC) using the Lorenz and Rössler system, where two unknown components are inferred from one given component. For both systems, NGRC achieves accurate results while requiring fewer training data and less computational time than RC. We identified an inverse proportional behavior between the number of time-delayed steps needed for NGRC and the temporal resolution, indicating that the physical time span covered by the delay interval is an important factor in determining the required number of delayed steps. Finally, we apply NGRC to the observational climate data of ENSO (El Niño--Southern Oscillation) and infer one observable from the remaining variables. Despite the noise and complexity of the real-world data, the NGRC shows promising results. Our findings demonstrate the potential of NGRC for efficient inference of unseen components in both controlled dynamical systems and real-world data.</description>
    </item>
  </channel>
</rss>
