Companies & labs
Straight from the people who built it
Every post here is an organisation publishing about its own work: a lab's newsroom, a research blog, a paper. No intermediary, no summary of a summary, and nothing this deck had to interpret. It is the shortest distance between you and the source.
01By company
- PrimaryAllen Institute for AI17 Sep, 3:00 AM CDT
What a crowdsourced game revealed about steering Olmo 3
A crowdsourced game built on Olmo 3 showed how people can exploit unexpected model behaviors to stress-test prosocial AI evaluations—and how open access to a model’s internals can help researchers understand why those tests break.
- PrimaryAllen Institute for AI14 Sep, 3:00 AM CDT
Teaching future scientists to interrogate AI tools for scientific discovery
University of Washington students put Ai2’s AutoDiscovery to the test, showing how AI can surface promising scientific leads while making human judgment, domain expertise, and rigorous validation more important than ever.
- PrimaryAllen Institute for AI9 Sep, 3:00 AM CDT
How Goodfire used Ai2’s open post-training stack to trace unwanted model behavior
Goodfire used Ai2’s fully open post-training stack to predict LLM behavioral changes, trace unwanted model behavior back to individual training examples, and test targeted fixes without sacrificing broader capability gains.
- PrimaryAllen Institute for AI1 Sep, 3:00 AM CDT
The hard parts of AI-assisted science
At an Ai2 event marking our expanded collaboration with Providence Swedish, researchers explored the hardest problems in AI-assisted science: keeping systems steerable, grounded in human judgment and sound methods, and responsive to new evidence and experiment
- PrimaryAllen Institute for AI1 Sep, 3:00 AM CDT
BenchMIRT: What are LLM benchmarks actually measuring?
BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and easier-to-interpret evaluations.
- PrimaryAllen Institute for AI27 Aug, 3:00 AM CDT
Ai2 and Providence Swedish Cancer Institute partner to advance AI-assisted scientific discovery
Ai2 and Providence Swedish Cancer Institute are expanding their collaboration after AutoDiscovery helped researchers uncover and validate a promising new immune signal in invasive lobular breast cancer.
- PrimaryAllen Institute for AI26 Aug, 3:00 AM CDT
How researchers adapted Dolma for better Thai language models
Thai researchers adapted Ai2’s open Dolma toolkit to build Mangosteen, a 47-billion-token Thai corpus that filters low-quality web data while maintaining or improving model performance and strengthening Thai cultural knowledge.
- PrimaryApple Machine Learning16 Sep, 7:00 PM CDT
REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying environmental reversibility, a property absent in real world manipulation, where events such as
- PrimaryApple Machine Learning15 Sep, 7:00 PM CDT
Shared Selective Persistent Memory for Agentic LLM Systems
Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use patterns that made previous sessions prod
- PrimaryApple Machine Learning15 Sep, 7:00 PM CDT
How Value Induction Reshapes LLM Behaviour
Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as helpfulness, harmlessness, and honesty. This is done to increase utility, ensure
- PrimaryApple Machine Learning15 Sep, 7:00 PM CDT
Enterprise data lakes accumulate tables faster than human stewards can document or classify them, leaving columns with missing descriptions and unassigned governance labels. This documentation debt undermines data discovery, access control, and regulatory comp
- PrimaryApple Machine Learning15 Sep, 7:00 PM CDT
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weak
- PrimaryApple Machine Learning10 Sep, 7:00 PM CDT
Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video descript
- PrimaryApple Machine Learning10 Sep, 7:00 PM CDT
DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language glo
- PrimaryarXiv21 Sep, 12:59 PM CDT
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either
- PrimaryarXiv21 Sep, 12:57 PM CDT
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory
- PrimaryarXiv21 Sep, 12:56 PM CDT
We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator locates the first inappropriate token and
- PrimaryarXiv21 Sep, 12:55 PM CDT
LoRA-generating hypernetworks for efficient on-device LLM generative personalization
On-device large language models (`LLMs'), e.g. running on mobile phones, are ripe for improvement via personalization. The limited compute resources of mobile devices impose limits on model scale and thus model quality, making any realizable quality gains high
- PrimaryarXiv21 Sep, 12:55 PM CDT
DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation
Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot dir
- PrimaryarXiv21 Sep, 12:55 PM CDT
Harness-Zero: Harness Distillation via Agent-as-Harness
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a ge
- PrimaryarXiv21 Sep, 12:54 PM CDT
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selec
- PrimaryarXiv21 Sep, 12:54 PM CDT
DolphinBench: Mapping the Pareto Frontier of Agent Memory
Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be ret
- PrimaryarXiv21 Sep, 12:53 PM CDT
Rare Event Estimation via Iterative Unalignment
As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We
- PrimaryarXiv21 Sep, 12:52 PM CDT
Emergent Collusion in Long-Horizon LLM Agent Interaction
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks,
- PrimaryarXiv21 Sep, 12:42 PM CDT
Learning Physics from an Imperfect Ancestor
Neural operators evaluate parametric partial differential equations cheaply but degrade sharply outside their training distribution. Physics-informed neural networks avoid dependence on labeled data, yet their optimization can be basin-fragile: when the govern
- PrimaryarXiv21 Sep, 12:39 PM CDT
Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization
A model generalizes outside its training distribution only when it computes a representation structurally equivalent to the generating mechanism, not an approximation fitted to it. Such equivalence is necessary for exactness in and out of distribution, and ext
- PrimaryarXiv21 Sep, 12:26 PM CDT
Linguistic Features for Interpretable Textual Entailment
Despite the success of neural models in natural language processing, their black-box nature limits interpretability and conceals the linguistic phenomena underlying their predictions. We present SLITE, an explainable hybrid model for Recognizing Textual Entail
- PrimaryarXiv21 Sep, 12:24 PM CDT
In this paper, we study nonasymptotic $L^p$ error bounds for interval length and conditional coverage in split conformalized quantile regression (CQR). Our bounds rely on local regularity conditions and accuracy guarantees for the estimated quantiles. We furth
- PrimaryarXiv21 Sep, 12:22 PM CDT
Et Tu, Brute? Economic Misalignment in Personal AI Agents
Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user's personal context, e.g., the
- PrimaryarXiv21 Sep, 12:18 PM CDT
BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction
Scientific weak signals are early, low-visibility research directions that later become central to mature scientific topics, yet existing resources such as trend tracking, citation forecasting, and foresight reports rarely provide validated reference sets that
- PrimaryarXiv21 Sep, 12:14 PM CDT
Social simulation offers the social sciences an experimental instrument that the real world cannot supply, and generative agents have transformed it by acting as silicon samples that unite agent-based modeling with real behavioral data. Existing platforms veri
- PrimaryarXiv21 Sep, 12:09 PM CDT
Visuomotor Robotic Pruning in Planar Orchards Using Hybrid Reinforcement Learning
Dormant tree pruning is labor-intensive yet essential for maintaining modern high-productivity fruit orchards. In this work, we focus on pruning of modern planar tree training systems - V-Trellis apples and UFO cherries - where trunks and primary branches are
- PrimaryarXiv21 Sep, 12:09 PM CDT
ToneCL: Contrastive Learning for Few-Shot Syllable-Level Tone Classification
Tone languages constitute over 50-70% of the world's languages, but the vast majority are low-resource, lacking the large transcribed corpora needed for automatic tone classification. Existing datasets are typically collected at the sentence level, whereas fie
- PrimaryarXiv21 Sep, 12:00 PM CDT
Human-LLM Deliberation as Interactive Proof: Conditions for Verifiability Without Transparency
When an LLM supplies an argument that a user could not readily construct, how can the user decide whether to accept its claim? Inspired by interactive proofs, we model human-LLM deliberation as an interaction between a prover with unrestricted internal search
- PrimaryarXiv21 Sep, 11:59 AM CDT
SLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models
Whole-slide pathology images (WSIs) contain gigapixel-scale visual content, creating a major scalability challenge for slide-level multimodal large language models (MLLMs). Existing approaches process thousands of patch tokens and typically apply compression o
- PrimaryarXiv cs.CL21 Sep, 11:55 AM CDT
OSWorld-Pro: Process-based Evaluation for Computer Use Agents
Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency i
- PrimaryarXiv cs.CL21 Sep, 11:53 AM CDT
The Copy Ceiling: An Input-Exposure Control for Ontology-Grounded Generation over Curated Corpora
When a language model answers from a curated corpus via graph-based retrieval, a large grounding uplift does not establish reasoning over the retrieved structure: the context may already expose the gold answers. We propose exposure accounting, which classifies
- PrimaryarXiv cs.CL21 Sep, 11:30 AM CDT
Argumentative component detection (ACD) is a core subtask of Argument(ation) Mining (AM) and one of its most challenging aspects, as it requires jointly delimiting argumentative spans and classifying them into components such as claims and premises. While rese
- PrimaryarXiv cs.CL21 Sep, 11:11 AM CDT
The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts
The Linear Representation Hypothesis associates high-level concepts with directions in language models, but it remains unclear how these concept-related linear structures are organized within the model. We propose the Answer-Basin Representation Hypothesis: th
- PrimaryarXiv cs.CL21 Sep, 11:04 AM CDT
MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents
Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, and collaborative work, are inherently multi-speaker. Supporting these settings introduces challenges that are
- PrimaryarXiv cs.CL21 Sep, 10:59 AM CDT
Post-training quantization (PTQ) enables efficient deployment of large language models, and PTQ methods are usually optimized and evaluated with generic reconstruction, perplexity, or answer accuracy. But in explanation-critical domains, preserving only the fi
- PrimaryarXiv cs.CL21 Sep, 9:47 AM CDT
Adapting Tree-Structured Speculative Decoding to DeepSeek-V4 for Efficient Inference
Repeated execution of the target model during autoregressive decoding is a major source of LLM inference latency. Unlike linear speculation, which follows a single candidate chain, tree-structured speculation retains multiple branches from shared prefixes; und
- PrimaryarXiv cs.CL21 Sep, 9:25 AM CDT
Circuit Hypernetworks for Quantum-Augmented Diffusion Language Models
Language models can be adapted by changing the computations applied to individual tokens. Quantum circuits offer one such approach, but evaluating wider circuits inside a large model can be computationally demanding. Here we introduce HyperQ, which adds token-
- PrimaryarXiv cs.CL21 Sep, 9:22 AM CDT
Assessing Readability with LLMs: The Role of Reasoning and Few-Shot Prompting
Readability assessment is essential for tailoring texts to intended audiences across educational, healthcare, and information retrieval domains. However, traditional readability formulas struggle to generalize across genres and languages, while supervised mach
- PrimaryarXiv cs.CL21 Sep, 9:16 AM CDT
Written as a Record, Read as an Address: What a Forward Pass Leaves in an Operation's KV Cache
When a language model reads an operation such as "Swap the contents of Box F and Box B", its forward pass writes keys and values for those tokens into the KV cache. Prior work on entity tracking establishes what models use: bindings are resolved at query time
- PrimaryarXiv cs.CL21 Sep, 9:09 AM CDT
Custom Named Entity Recognition and Topic Classification for Global Health Publications
How should natural language processing models be selected and adapted for global health literature in environments where annotated data and computational resources are limited? This thesis investigates these challenges through experiments on semantic tag disco
- PrimaryarXiv cs.LG21 Sep, 11:50 AM CDT
Learning Prognostic Variables for AI Convective Parameterizations via Symbolic Distillation
Hybrid AI-physics climate modeling aims to improve coarse (~100km-resolution) Earth system models by learning to parameterize subgrid processes from high-fidelity data. However, this so far mostly involves local-in-time, diagnostic parameterizations, in which
- PrimaryarXiv cs.LG21 Sep, 11:33 AM CDT
When Tomorrow Becomes Today: Self-Evolving Policies for Agentic Time-Series Forecasting
Agentic time series forecasting concerns systems whose underlying mechanisms evolve, making the relative effectiveness of numerical models, reasoning strategies, and intervention rules inherently time-varying. Consequently, a time series agent must adapt the f
- PrimaryarXiv cs.LG21 Sep, 11:20 AM CDT
PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control
Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker's capabilities, leaving r
- PrimaryarXiv cs.LG21 Sep, 11:11 AM CDT
G-NAC: Graph Neural Automata Clustering via Emergent Domain Formation
We introduce Graph Neural Automata Clustering (G-NAC), an unsupervised clustering method in which observations interact as cells on a fixed neighborhood graph. A shared recurrent graph-neural cellular rule evolves latent domain states through local interaction
- PrimaryarXiv cs.LG21 Sep, 10:50 AM CDT
Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, voc
- PrimaryarXiv cs.LG21 Sep, 10:38 AM CDT
Localising artifacts in synthetic speech remains challenging, as most evaluation methods yield only global quality scores. This paper presents XSQ-AST, a framework that combines the SQ-AST speech quality model with WhisperX phoneme alignment and multiple salie
- PrimaryarXiv cs.LG21 Sep, 10:21 AM CDT
We investigate next generation reservoir computing (NGRC) as a data-driven approach for inferring unseen components of dynamical systems. We compare NGRC with traditional reservoir computing (RC) using the Lorenz and Rössler system, where two unknown component
- PrimaryarXiv cs.LG21 Sep, 10:17 AM CDT
Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
The growing demand for real-time, data-driven decision-making in complex and dynamic systems is placing increasing pressure on traditional Operational Research (OR) methodologies. Reinforcement learning (RL) has emerged as a complementary approach, offering st
- PrimaryarXiv cs.LG21 Sep, 10:17 AM CDT
D-JEPA: A Decision-Aligned Latent World Model
Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execut
- PrimaryarXiv cs.LG21 Sep, 10:16 AM CDT
Enhancing Transformer Representations of Symbolic ODE Expressions
Existing approaches to solving differential equations, such as symbolic regression, physics informed neural networks, and neural operators, typically focus on numerical approximations or blind symbolic search via fitting to numerical data. Less attention has b
- PrimaryarXiv cs.LG21 Sep, 9:56 AM CDT
While a centralized approach involving patient consent to collect and analyze data centrally would theoretically offer the best data quality and predictive performance, it is not always feasible in practice. Federated Learning (FL) architectures have shown to
- PrimaryAWS Machine Learning22 Sep, 1:10 PM CDT
Bring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock
GPT-6 Sol and GPT-6 Luna are now generally available on Amazon Bedrock, giving you more options to match intelligence and efficiency to each workload.
- PrimaryAWS Machine Learning22 Sep, 12:28 PM CDT
Claude Opus 5.5 is now available on AWS
Claude Opus 5.5, Anthropic's most capable Opus model for agentic coding, knowledge work, and long-running tasks, is now available on Amazon Bedrock and Claude Platform on AWS. This post covers what's new in Opus 5.5, practical guidance, and how to start buildi
- PrimaryAWS Machine Learning22 Sep, 12:18 PM CDT
Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore
Skills let you encode domain-specific procedures as reusable, portable instructions for agents, but a fluent answer doesn't prove the agent picked the right skill or followed it. Learn how to measure skill selection and instruction following with Strands Evals
- PrimaryAWS Machine Learning22 Sep, 10:46 AM CDT
How Reactiv automates mobile commerce 80% faster with Amazon Bedrock AgentCore
Reactiv used Amazon Bedrock AgentCore to build a multi-agent AI Scheduler that autonomously refreshes Shopify merchants' mobile apps on a schedule, reducing merchant configuration time by 80% and getting to production 33% faster.
- PrimaryAWS Machine Learning22 Sep, 10:35 AM CDT
Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI
Concurrency sweeps help you right-size a generative AI endpoint on Amazon SageMaker AI by systematically benchmarking it at increasing load levels. This post walks through deploying a model, running automated concurrency sweeps with the CreateAIBenchmarkJob AP
- PrimaryAWS Machine Learning22 Sep, 10:30 AM CDT
How Trane gets building insights 60x faster with Amazon Bedrock AgentCore
In about four weeks, Trane Technologies built an AI-powered agentic solution on Amazon Bedrock AgentCore that reduced a 20-minute, multi-screen building diagnostic workflow to a 20-second natural language interaction, a 60x improvement in time-to-insight. This
- PrimaryAWS Machine Learning22 Sep, 10:19 AM CDT
How Tata Elxsi detects industrial safety risks in seconds on AWS
Learn how Tata Elxsi built IRIS, a real-time industrial safety platform on AWS. IRIS filters camera video at the edge, streams metadata through Amazon Kinesis, runs computer vision on Amazon SageMaker AI, and correlates detections into high-confidence alerts,
- PrimaryAWS Machine Learning22 Sep, 10:17 AM CDT
Extending public sector intelligence with Agentforce and AWS
Public sector agencies process large volumes of unstructured evidence, such as body camera footage and scanned documents. This post shows how to combine Amazon Bedrock Data Automation with the Model Context Protocol (MCP) to turn that data into structured insi
- PrimaryAWS Machine Learning21 Sep, 1:30 PM CDT
xAI’s Grok 4.6 is now available in Amazon Bedrock
xAI's Grok 4.6 is now available in Amazon Bedrock: a frontier model for long-running agents, coding, and knowledge work, with a 500K token context window and four reasoning effort levels. It runs on both the bedrock-mantle and bedrock-runtime endpoints, with C
- PrimaryAWS Machine Learning21 Sep, 11:34 AM CDT
Run Positron on Amazon SageMaker AI for data science workflows
Positron, Posit's IDE for data science, now runs on Amazon SageMaker AI. This post shows how a data scientist explores an Amazon Athena table, validates features in R, trains an XGBoost model in Python, deploys a real-time SageMaker AI endpoint, and reports re
- PrimaryAWS Machine Learning21 Sep, 11:27 AM CDT
How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore
Learn how Benchling built a defense-in-depth security architecture to run untrusted, AI agent-generated scientific code across thousands of life sciences tenants using Amazon Bedrock AgentCore Code Interpreter in VPC mode, combined with Amazon Route 53 Resolve
- PrimaryAWS Machine Learning21 Sep, 11:24 AM CDT
Reducing medical claims review time with AI on AWS: The EXL Medical IDP solution
EXL built an AI-powered Medical intelligent document processing (IDP) solution on AWS, combining IDP with domain-specific large language models on Amazon SageMaker and Amazon Bedrock to extract, summarize, and query medical records at enterprise scale and cut
- PrimaryAWS Machine Learning18 Sep, 3:52 PM CDT
Amazon SageMaker Inference: 2026 year-to-date launches in review
Amazon SageMaker AI shipped 13 inference launches in year-to-date across two deployment paths: fully managed endpoints and Amazon SageMaker HyperPod Inference. This post reviews each launch, from inference recommendations and capacity-aware instance pools to t
- PrimaryAWS Machine Learning18 Sep, 11:52 AM CDT
Introducing Kimi K3 on Amazon Bedrock
Kimi K3 from Moonshot AI is now available on Amazon Bedrock, giving you a powerful new open-weight option for coding and knowledge work. It offers native vision, a 1-million-token context window, and explicit prompt caching to reduce latency and input costs.
- PrimaryAWS Machine Learning18 Sep, 10:38 AM CDT
Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime
Migrate a multi-model healthcare AI agent from self-managed Amazon ECS with AWS Fargate to Amazon Bedrock AgentCore runtime, preserving triple-model orchestration and vector-enhanced knowledge retrieval while reducing infrastructure management. The framework-a
- PrimaryAWS Machine Learning18 Sep, 10:31 AM CDT
The new AgentCore runtime: Elastic, optimized, and consistently fast starts
Today we are announcing the new AgentCore runtime, a capability of Amazon Bedrock AgentCore built for the speed, flexibility, and cost efficiency that production agents demand. It reclaims memory as sessions release it and delivers consistent cold starts regar
- PrimaryAWS Machine Learning18 Sep, 10:25 AM CDT
Deploy Hugging Face models on Amazon SageMaker AI with coding agents
Deploy production-ready Hugging Face models on Amazon SageMaker AI using six open-source agent skills. Point a coding agent at a model and get back a real-time endpoint with the right serving container, autoscaling, Amazon CloudWatch alarms, and a verified tea
- PrimaryAWS Machine Learning18 Sep, 8:08 AM CDT
Introducing Amazon SageMaker HyperPod Inference Gateway
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing add-on for Amazon EKS. It uses real-time GPU signals to send each inference request to the best-suited pod, cutting first-token latency by up to 82% with no changes to your m
- PrimaryAWS Machine Learning17 Sep, 12:55 PM CDT
Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent
Amazon Connect Talent is an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, and consistent evaluation, helping recruiters identify strong candidates more efficiently while
- PrimaryBerkeley BAIR29 Jul, 4:00 AM CDT
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new epoch in computing. Hardware is changing rapidly — not just
- PrimaryBerkeley BAIR26 Jul, 4:00 AM CDT
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Overview of ABBEL compared to traditional recursive summarization. Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performance by supervising the contents of each belief state.. As task horizons grow, LL
- PrimaryBerkeley BAIR7 Jul, 4:00 AM CDT
Intelligence is Free, Now What? Data Systems for, of, and by Agents
... government of the people, by the people, for the people ... , Abraham Lincoln, Gettysburg Address (1863) The cost of AI is dropping rapidly. GPT-4-class capabilities cost roughly $30 per million tokens in early 2023; today the same runs under $1 , and some
- PrimaryBerkeley BAIR1 Jul, 4:00 AM CDT
Congratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence
02How a post gets here
Evidence tier 1, and nothing else. This deck defines that tier as the organisation said it itself, or this is the paper, and the tier is set per source from what the source IS rather than from how good it is. So a lab's own newsroom qualifies and an outlet reporting on that lab does not, however good the reporting. 21 of the deck's sources carry that tier.
Every link went through the reader's door first, the same as every article link on this site: 0 post(s) were dropped because a reader could not open them. And this page has its own feed at companies.xml, carrying the same items under the same permission law as the main feed.
Publishing here: AWS Machine Learning · Allen Institute for AI · Apple Machine Learning · Berkeley BAIR · EleutherAI · Google AI · Google DeepMind · Google Research · Hugging Face · IBM Research · Microsoft Research · Mistral AI · NIST · NVIDIA · OpenAI · Qwen (Alibaba) · Stability AI · Together AI · arXiv · arXiv cs.CL · arXiv cs.LG