FuguReport

FuguReport

Browse the latest weekly themes first, then scan the most recent daily reports and archives.

Anchor Date: 2026-08-27
Weekly

2026-08-14 - 2026-08-20

Daily

Recent Daily Reports

50 reports
2026-08-27 Task / Surface Detection / Glass surface detection in scenes

Glass Surface Detection Grounded in 3D Visual Geometry

This paper reframes glass surface detection as a 3D visual geometry-grounded task rather than relying predominantly on 2D appearance cues.

2026-08-27 Evaluation / Benchmarking / Probabilistic alignment evaluation

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

This paper argues that evaluating video generators as visual world models requires assessing the distribution of possible futures under a fixed observation and action, rather than judging only single-rollout plausibility.

2026-08-27 Method / Reinforcement Learning / Policy optimization via rollout agreement

TTPO: Test-Time Policy Optimization

This paper introduces Test-Time Policy Optimization (TTPO), a label-free test-time training method for mathematical reasoning models.

2026-08-27 Task / Backdoor Injection / Trapping failure modes in VLA models

TrapVLA: Trapping Vision-Language-Action Models in Configured Failure Modes

This paper studies a new backdoor attack setting for Vision-Language-Action models called Configured Failure Trapping, where a stealthy textual trigger causes a robot to fail in a specific, preconfigured way rather than simply fail arbitrarily.

2026-08-27 Method / Self-Distillation / On-policy distillation framework

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

This paper presents Self-OPD, a teacher-free on-policy distillation framework for flow matching models.

2026-08-26 Method / Memory Networks / Dual-brain memory architecture

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

VoiceMem is a memory architecture for real-time spoken interaction that separates memory into a factual "left brain" and an affective/persona-oriented "right brain," connected through streaming memory I/O.

2026-08-26 Method / Agent / Task-adaptive agent harness synthesis

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

This paper presents JIT-Agent, a model trained to generate task-adaptive agent harnesses on demand rather than relying on a fixed, manually engineered scaffold.

2026-08-26 Evaluation / Instruction Following Evaluation / Multimodal LLMs on video tasks

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Video-IFBench is a benchmark designed to evaluate how well multimodal large language models follow instructions in video understanding settings, rather than only measuring answer correctness.

2026-08-26 Method / Reinforcement Learning / Visuals-based reinforcement learning approach

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

This paper studies visual faithfulness in vision-language model post-training and frames the problem as a credit-assignment issue rather than only an evaluation issue.

2026-08-26 Method / Video Action Modeling / In-context causal modeling from videos

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

This paper frames zero-shot cross-task robotic manipulation as an in-context learning problem, where an unseen task is specified at deployment time by a human demonstration video rather than only by language.

2026-08-25 Method / Model Safety / Step-level tool action checking

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

StepGuard is a 4B step-level guard model for LLM agents that supports both pre-execution checking of candidate tool actions and post-hoc auditing of completed trajectories.

2026-08-25 Application / Dialogue Systems / Interactive dialogue game tasks

Investigating Knowledge Transfer Across Interactive Dialogue Games

This paper studies how knowledge transfers across interactive dialogue games by fine-tuning a common LLM on tasks from the clembench suite and evaluating cross-task performance.

2026-08-25 Method / Reinforcement Learning / Segment-level credit assignment techniques

Ockhamareto: Pareto-Gated Segment-Level Credit Assignment for Concise Unit-Test Generation with Reinforcement Learning

Ockhamareto is a single-shot GRPO framework for generating complete unit-test suites while jointly optimizing fault-detection effectiveness and suite conciseness.

2026-08-25 Method / Speech Synthesis / Synthetic data generation for medical dialogues

Speech-to-SOAP: End-to-End Summarization of Medical Dialogues: KIT@BeTraC 2026

This paper presents KIT's BeTraC 2026 lightweight-track system for generating SOAP clinical notes directly from medical dialogue audio.

2026-08-24 Method / Multi-Agent Systems / Meta-harness optimized agent framework

MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters

This paper studies few-shot learning for lightweight time series forecasters, targeting resource-constrained settings where large foundation models are too costly and only small support sets are available.

2026-08-24 Evaluation / Scientific Agent Evaluation / Benchmarking scientific agents on Earth systems

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

EarthVerse is a benchmark designed to evaluate scientific agents conducting multi-source investigations across dynamic Earth systems and natural hazards.

2026-08-24 Method / World Modeling / Interactive and high-fidelity multi-step modeling

ReWorld: An Interactive World Model with Long-Horizon Memory

ReWorld is an interactive streaming world model designed to simultaneously follow user camera actions, preserve memory of previously visited places, and run in real time.

2026-08-24 Method / Reinforcement Learning / Scaling RL for diffusion models

Scaling Reinforcement Learning for Diffusion Models via Velocity Matching

This paper argues that reward fine-tuning for diffusion models does not need likelihood-based policy-gradient machinery and instead can be formulated directly in the model's native velocity representation.

2026-08-24 Method / Optimization / Iterative harness updates from failures

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

This paper presents AutoSaddler, a framework for automatically improving the external harness of LLM agents on long-horizon tasks.

2026-08-23 Method / Motion Prediction / Two-stage framework with intermediate representation

EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting

EMPIRE is a two-stage framework for egocentric bimanual hand-motion forecasting that inserts an explicit manipulation plan between multimodal observation and motion generation.

2026-08-23 Method / MoE Routing / Routing-guided steering and selection

Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents

This paper studies whether native mixture-of-experts (MoE) router traces can guide test-time scaling for software-engineering agents without using an external judge or running candidate patches for selection.

2026-08-23 Method / Skill Reliability / Reliability interventions in skill selection

Coalition-Aware Skill Reliability for Self-Evolving Agents

This paper investigates whether skills stored in self-evolving LLM agents' skill banks make positive mechanistic contributions rather than merely improving aggregate outcome metrics.

2026-08-23 Evaluation / Model Safety Evaluation / Robustness against rebuttal attacks

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

This paper studies a post-decision attack surface in LLM-based hate speech moderation: fabricated annotator-style rebuttals that ask a model to reconsider an initially correct judgment.

2026-08-23 Method / Reinforcement Learning / Stage-aware dialogue modeling

DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue

DeepSAGE is a hybrid LLM–DRL framework for conducting the first CBT counseling session as an explicit sequence of eleven stages, each with a therapeutic objective and automated stage-completion criterion.

2026-08-22 Method / Network Security / ML-based intrusion prevention

Autonomous Cyber Defense: Real-Time Attack Detection and Mitigation in Software-Defined Networks Using Machine Learning

This paper presents an automated cyber defense system for software-defined networks that combines real-time traffic monitoring, machine-learning-based attack diagnosis, and automatic mitigation.

2026-08-22 Evaluation / Benchmarking / Benchmark for SfM in real-world scenarios

ORBIT++: Benchmarking SfM in the Wild with 360° Video

This paper presents ORBIT++, an updated version of the ORBIT benchmark for evaluating structure-from-motion and camera pose estimation on difficult real-world videos derived from online 360° footage.

2026-08-22 Method / Ranking Metrics / Normalized discounted cumulative gain reformulation

Revisiting N2DCG: An Empirically Grounded Reformulation of Carousel Recommendation Evaluation

This paper re-examines N2DCG, a two-dimensional adaptation of NDCG for carousel recommendation interfaces, and argues that the original formulation makes assumptions that do not hold in real carousel layouts.

2026-08-22 Evaluation / Benchmarking / Composable compression quality assessment

Benchmarking Composable Compression Techniques in Mixture-of-Experts LLMs

This paper introduces MoE-XBench, a benchmark for evaluating compression in mixture-of-experts large language models as an end-to-end deployment workflow rather than as isolated steps.

2026-08-22 Method / State Management / Techniques for storing utterance context in ASR

Beyond Fresh Starts: Stateful Inference for Streaming ASR in Conversational Voice Agents

This paper analyzes how state handling in low-latency streaming ASR affects conversational voice agents, focusing on turn boundaries, long silences, and short backchannel turns.

2026-08-21 Application / Multi-Agent Coordination / System intelligence with LLM agents

Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence

This paper provides a comprehensive survey on how LLM-based systems are transitioning from single-agent individual intelligence to multi-component system intelligence to tackle complex, long-horizon tasks.

2026-08-21 Method / Adversarial Attacks / Critic-induced value-subspace attacks

CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents

This paper studies white-box, causal, online adversarial attacks against visual world-model agents, focusing on a frozen DreamerV3 victim.

2026-08-21 Method / Federated Learning / Personalized FL with Lorentz embeddings

FlatLand: Personalized Graph Federated Learning via Tailored Lorentz Space

FlatLand is a personalized federated learning method for graph data that places each client in a tailored Lorentz-space embedding rather than assuming a shared Euclidean geometry.

2026-08-21 Method / Natural Language Generation / Generating adaptive explanations for errors

Beyond the Traceback: Using LLMs for Adaptive Explanations of Programming Errors

This paper studies whether LLM-rewritten Python error messages help programmers debug more effectively than standard interpreter output.

2026-08-21 Method / Policy Learning / Physics-informed code-as-policy agent

PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration

PhysCaP is a physics-informed code-as-policy agent designed for active perception in robotic manipulation tasks where success depends on latent physical properties that cannot be inferred from vision alone.

2026-08-20 Evaluation / Visual Reasoning Evaluation / Assessing visual inference in video models

VGI-BENCH: Probing Visual Intelligence in Video Generation Models

This paper introduces VGI-Bench, a benchmark for evaluating visual intelligence in video generation models through process-sensitive, visually grounded tasks.

2026-08-20 Method / Optimal Transport / Coupled OT framework with landmarks

Coupled Optimal Transport with Landmark Constraints

This paper addresses a limitation of standard optimal transport: minimizing transport cost alone may fail to identify a geometrically meaningful transformation between distributions.

2026-08-20 Method / Model Routing / Query routing to experts

Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

This paper studies AI model routing when estimating each candidate specialist’s value is itself costly.

2026-08-20 Evaluation / Subjectivity Evaluation / Impact of subjectivity on accuracy and training

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

This paper argues that standard accuracy-based evaluation and hard-label training are poorly matched to LLM-based social simulation because human responses are inherently distributional rather than deterministic.

2026-08-20 Method / Self-Supervised Learning / Framework for audio representation

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

This paper introduces NAPE, a self-supervised audio learning framework in which a causal Transformer predicts the embedding of the next log-mel spectrogram patch from preceding patches.

2026-08-19 Method / Graph Neural Networks / Temporal-spatial particle graph transformer

DyG$^2$T: Modeling Object Dynamics with 3D Gaussian Temporal-Spatial Particle Graph Transformer

This paper presents DyG2T, a framework for predicting object motion from limited visual observations by combining dynamic 3D Gaussian reconstruction with dynamics modeling over downsampled key points.

2026-08-19 Method / Reward Modeling / Chain of omni reward model for multimodal generation

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

This paper studies reward modeling for joint video-audio generation using human preference feedback rather than separate automatic metrics.

2026-08-19 Evaluation / Capability Analysis / Empirical study of agent abilities

What is Missing from AI Post-Training AI: An Empirical Analysis

This paper investigates AI agents performing end-to-end post-training of large language models, distinguishing between execution-level capability (iterating within an established pipeline) and strategy-level capability (revising high-level paradigms and workflows based on evidence).

2026-08-19 Method / Agent / Unified agentic system for documents

DocClaw: A Unified Agentic System for Intelligent Document Processing

DocClaw is a unified agentic system for intelligent document processing that frames OCR, document question answering, and key information extraction as iterative interactions between an agent and a document.

2026-08-19 Method / Skill Learning / Training in-policy skill selection

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

This paper studies in-policy skill selection for long-horizon agents that must choose which skill file to read from a candidate slate during an episode.

2026-08-18 Evaluation / Benchmarking / Multi-domain AI system evaluation

ASI-Bench: At the Dawn of Artificial Superintelligence

ASI-Bench is a benchmark for evaluating whether AI systems can conduct project-level scientific research as human methodological guidance is progressively reduced.

2026-08-18 Evaluation / Agent Safety Evaluation / Lifecycle phase safety benchmarking

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

This paper introduces HarnessRisk, a lifecycle-oriented benchmark for agent harness safety that organizes risks into six operational phases: Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery.

2026-08-18 Method / Interpretable Models / Internal readout for reasoning states

Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing

This paper introduces a two-level internal readout for mixture-of-experts reasoning models that aims to expose latent reasoning state beyond what is written in the generated trace.

2026-08-18 Method / Visual Prompting / Compositional prompt learning framework

Primitive-Driven Compositional Forensic Visual Prompting for Open-World Face Anti-Spoofing

This paper studies open-world face anti-spoofing, where models must handle both cross-domain shifts and previously unseen attack types.

2026-08-18 Method / Recommendation Models / LLM-based ranking framework

Empowering Compact LLMs with Fusion of Layer-wise Exits for Recommendation

This paper studies scalable sequential recommendation with compact large language models under a discriminative, embedding-based ranking setup.

2026-08-17 Method / Video Editing / Visual in-context video editing paradigm

VicEdit: Learning to Edit Videos from Visual In-Context Examples

This paper introduces Visual In-context Editing, a video editing paradigm that supplements text instructions with visual examples in the form of a single image, an image pair, or a video pair.

Anchor Date: 2026-08-27
Archive

Archive

Weekly Archive

15

インタラクティブモデルとエージェントのためのベンチマーク

This week's theme centers on moving model evaluation beyond narrow offline metrics toward executable, task-grounded benchmarks for embodied world models and multimodal GUI agents.

2026-08-14 - 2026-08-20

幾何学を考慮した3D再構成と生成

This theme concerns making visual generation and reconstruction explicitly grounded in geometry, camera parameters, and scene structure.

2026-08-14 - 2026-08-20

スケーラブルな汎用強化学習

This week's reinforcement learning work focuses on making RL systems more general, computationally efficient, and stable under demanding training regimes.

2026-08-14 - 2026-08-20

科学LLMの評価と適応

This theme centers on how foundation models are being evaluated and improved for scientific and multimodal use, with emphasis on pretraining strategy, data realism, and efficient cross-modal design.

2026-08-07 - 2026-08-13

スパースニューラル表面再構成

This week's theme centers on making neural surface reconstruction both generalizable and efficient under sparse-view conditions.

2026-08-07 - 2026-08-13

自己進化型LLMマルチエージェントシステム

This theme tracks the shift from static, centrally orchestrated LLM agents toward self-evolving, decentralized multi-agent systems that adapt their roles, memory, tools, and coordination over time.

2026-08-07 - 2026-08-13

根拠付きマルチモーダルQAベンチマーク

Benchmarking work is increasingly targeting harder, more credible evaluation settings for multimodal question answering, especially where models must retrieve or ground evidence rather than rely on shallow multiple-choice cues.

2026-07-31 - 2026-08-06

コンピュータ操作エージェントの評価

Evaluation for computer-use and GUI agents is shifting from task success measurement toward diagnosis of safety, efficiency, and cross-platform capability.

2026-07-31 - 2026-08-06

オンライン模倣学習と世界モデル強化学習

Representative papers converge on a shared limitation of static imitation and standard model-based control: policies struggle with distribution shift, sparse rewards, and weak exploration once they leave the expert data regime.

2026-07-31 - 2026-08-06

マルチモーダル推論評価

This theme centers on how increasingly capable language and multimodal models should be evaluated and controlled as they move beyond text-only benchmarks.

2026-07-24 - 2026-07-30

マニピュレーションの汎化と評価

This week's theme centers on evaluating and improving embodied manipulation in contact-rich settings where vision-only supervision and narrowly collected robot data are insufficient.

2026-07-24 - 2026-07-30

LLMソフトウェア工学評価

This theme centers on evaluating AI coding assistants and software-engineering agents beyond narrow benchmarks or final-output metrics.

2026-07-24 - 2026-07-30

LLMコードエージェントの評価

This theme centers on evaluating LLM-based coding agents on realistic software-engineering tasks, especially repository-level issue resolution.

2026-07-17 - 2026-07-23

ロボット学習のための身体性世界モデル

This week's theme centers on using embodied world models not just to generate realistic futures, but to evaluate, supervise, and improve robotic policies.

2026-07-17 - 2026-07-23

対話型LLM行動評価

This week's theme centers on evaluating LLM behavior in interactive, socially grounded settings rather than judging single-turn text quality alone.

2026-07-17 - 2026-07-23
This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.