FuguReport
Browse the latest weekly themes first, then scan the most recent daily reports and archives.
2026-09-18 - 2026-09-24
Memory and Representations for Autoregressive Video
Work on autoregressive video models centers on two linked problems: preserving useful visual history over long rollouts and learning representations that keep generation efficient and coherent.
Theme 2Open LLM Post-Training and Evaluation
The theme centers on open, multi-stage LLM post-training in which evaluation and verification guide data selection, preference tuning, and reinforcement learning.
Theme 3Adaptive Multimodal Representation Learning
The theme centers on adapting representations to heterogeneous visual and linguistic inputs rather than relying on uniform processing.
Recent Daily Reports
From Source Code to Network Profile: Automated and Traceable MUD Profile Generation for IoT Devices
The paper presents AutoMUD, a source-code-driven system for generating Manufacturer Usage Description (MUD) profiles for IoT devices from firmware and software source trees.
Strategically Diverse Sampling for Self-Training
This paper investigates self-training data construction for large language models, arguing that repeated samples should vary in problem-solving approach rather than merely in surface form or correctness.
Online Learning via Learned Latent Bayesian Tracking
The paper studies online learning in non-stationary environments where models must adapt from streaming labeled samples under tight computational constraints.
GraphWrit3R: End-to-End 3D Scene Graph Writing
GraphWrit3R is an end-to-end method for generating 3D scene graphs directly from a point cloud, Gaussian splats, or both, outputting a structured JSON graph of objects, attributes, and relationships.
Recommendation World Models for Future-State Control
This paper studies how a trained sequential recommender can be wrapped with a decision layer that reasons about the future consequences of its displayed slates.
Parts-of-Speech as Emergent Categories in SAE Latent Space
This paper studies how part-of-speech (PoS) information is represented in Sparse AutoEncoder (SAE) latent activations extracted from LLaMA-3-8B.
Likelihood Ranking doesn't Scale Like Prompting in LLMs
This paper compares two evaluation protocols for multiple-choice question answering in large language models: direct prompted answering and likelihood ranking of declarative statements derived from the same question-answer pairs.
ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
ExplorationBench is a benchmark for evaluating whether AI systems can acquire new knowledge through exploration rather than recall.
Hyperbolic Multimodal Continual Learning: A Closest-Admissible Solution
This paper studies continual learning for hyperbolic multimodal models, where sequential updates can preserve task scores yet still distort the Lorentz geometry that encodes within-modality similarity, cross-modal alignment, and semantic hierarchy.
OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction
OmniFabric is a method for generating simulation-ready garment textures from a single in-the-wild clothing image by synthesizing textures directly in 2D sewing-pattern UV space.
Hunyuan-A13B Technical Report
Hunyuan-A13B is an open-source Mixture-of-Experts language model from Tencent featuring 80 billion total parameters while activating only 13 billion parameters per token.
Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents
This paper proposes Motion Vision CAPTCHA (MVCAP), a hierarchical CAPTCHA framework that hides target semantics in motion-defined foreground structures that become recoverable only through temporal segregation from a dynamically evolving background.
Safety-Aware Zero Trust Enforcement for IoT and Cyber-Physical Systems
This paper studies how Zero Trust enforcement should be adapted for IoT and cyber-physical systems, where restricting a suspicious component can reduce cyber risk but also remove sensing or control capabilities needed for safe operation.
InternW0: A Foundational Physical World Model for Efficient Real-World Interactions
InternW0 is the first model in the InternW physical world model series, designed to connect future visual prediction with continuous robot control for closed-loop interaction in the physical world.
Planned Test-Time Scaling with Coordinated Reasoning Paths
This paper proposes Planned Test-Time Scaling (PTTS), a framework for reasoning-time scaling that coordinates multiple inference branches instead of sampling them independently.
Risk-Aware Online Conformal State Probing
The paper studies a sequential decision-making setting in which a centralized agent must choose both control actions and whether to probe the true state from a robot or edge device.
Matryoshka attribution: Learning to attribute language model outputs to representations and weights
This paper frames attribution in language models as learning an ordering over nested subsets of internal components that minimizes a downstream loss.
Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces
This paper introduces Ladders-of-Thought (LoT), a training framework for improving reasoning in small- and mid-scale language models by combining progressive question rewrites with an adaptive curriculum.
From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs
This paper studies parameter-efficient fine-tuning for mixture-of-experts (MoE) language models and argues that expert-level adaptation is still too coarse because activated experts are internally sparse.
RoboTwin-Phys: Do WAMs and VLAs Understand the Physical World?
RoboTwin-Phys introduces physical-condition diversity as an explicit evaluation axis for robot manipulation, extending the 50-task RoboTwin-2.
Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
This paper analyzes the expressivity limits of Kimi Delta Attention (KDA) and introduces Complex KDA (CKDA), a variant that extends the gate range to [-1,1] and the delta-rule coefficient to [0,2].
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
WorldCrafter is a camera-controllable autoregressive video world model designed to preserve scene content over long horizons and across viewpoint changes.
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
onPanda is an interactive annotation tool for LLM alignment data and agent trajectories built around token-level correction.
MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes
This paper introduces a probe-based framework for evaluating vulnerability reports by replaying an agent's exploit and running executable checks of application security properties.
GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
This paper proposes the Geometry-Native Autoencoder (GAE), a compact latent space designed as a shared foundation for both 3D perception and generative visual modeling.
Accurate Motion Estimation with Bézier Control Point for Efficient Frame Interpolation
This paper presents ABC-Inter, a forward-warping video frame interpolation method designed to address motion ambiguity in training data and the common inference-time assumption of uniform linear motion.
Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment
This paper argues that alignment to human preferences is not the same as alignment to human behavior, and formalizes the discrepancy as a "Turing-test gap.
MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs
This paper introduces MinCU, a benchmark for grounded minimal-change understanding in near-identical image pairs.
VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks
VibeMemBench is a benchmark designed to evaluate persistent memory systems in coding agents on real repository tasks with executable validation.
MoSAT: Human Motion Generation from Spatial Audio and Textual Description
This paper studies human motion generation jointly conditioned on spatial audio and natural-language descriptions, arguing that the two modalities provide complementary information about what action to perform and where and when it should unfold.
Optimizers for Diffusion Models: A Controlled Benchmark
This paper presents a controlled benchmark of seven optimizers for diffusion training across four formulations: masked diffusion on text8, uniform diffusion on QM9, DUO on LM1B, and Gaussian image diffusion on CelebA-64.
PanoSeg3R: Feed-Forward 3D Semantic Segmentation for Panoramic Images with an Automatic Data Curation Pipeline
PanoSeg3R is a feed-forward framework for 3D semantic segmentation from multi-view panoramic equirectangular images.
SomaNet: Weakly Supervised Learning for Instance Soma Segmentation in 3D Electron Microscopy with Partial Annotations
SomaNet is a weakly supervised framework for 3D electron microscopy soma instance segmentation designed for settings with partial instance annotations.
Visual Navigation Transformer with Pose Attention
This paper proposes VNT-PA, a visual navigation transformer that represents an environment as a set of depth keyframes indexed by camera pose rather than as a temporally ordered observation history or an explicit map.
CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
CogGym is a scalable framework for comparing AI models and humans on matched cognitive-science experiments rather than standard answer-key benchmarks.
BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings
BrainWideBench is a benchmark for evaluating across-animal transfer in multi-region neural recordings, built on the International Brain Laboratory Brainwide Map dataset.
Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering
This paper studies cross-domain visual question answering for printed circuit board assembly (PCBA), where models must reason over fine-grained visual evidence, component semantics, and manufacturing standards across both standards-derived and real-world production images.
Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees
This paper introduces Conformal Privacy Auditing (CPA), a release-time privacy auditing framework for text that quantifies document-level re-identification risk under a declared attacker model.
ClashBench: Conflicts Leading Agents to Seize and Harm
This paper identifies and formalizes a safety failure mode in privileged agent systems called destructive resource preemption, where an agent resolves a resource conflict by disrupting an incumbent task instead of reporting the conflict or deferring to the user.
What Does Privileged Information Add to On-Policy Self-Distillation?
This paper investigates the degree to which privileged information (PI) contributes to on-policy self-distillation (OPSD) beyond the baseline gains of cross-mode distillation itself.
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
This paper introduces Video DeltaNet (VDN), a hybrid attention architecture for long-sequence video generation that keeps exact Softmax attention for local temporal interactions and boundary-anchor frames while using bidirectional linear memory for distant context.
Local Sparsity Enables Unsupervised LLM Safety Detection
This paper studies deployment-time LLM safety detection without unsafe training labels by framing the task as one-class anomaly detection on model activations.
Towards Active Cross-View Object Geo-Localization
This paper introduces ActiveGeo, a reformulation of cross-view object geo-localization in which a mobile agent sequentially acquires additional viewpoints and decides when to stop, rather than relying on a single fixed query image.
From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning
This survey examines Efficient Multimodal Learning (EML) as a full-stack problem spanning model design, algorithmic compression, and system deployment.
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
This paper analyzes a failure mode in PPO critic learning for reinforcement learning with large language models, termed Value Flattening.
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model
This paper evaluates whether the omni-modal generative model MiniMax-H3 can reason about the physical world when key information is distributed across multiple input modalities rather than fully specified in text prompts.
Reinforcement Learning for Real-Time Vision-Language-Action Policies
This paper addresses reinforcement learning fine-tuning for large vision-language-action (VLA) policies operating under non-negligible inference latency, which otherwise causes distribution shift from stale observations.
GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models
This paper introduces GYROval, a benchmark for measuring cultural value orientation in large language models along the two Inglehart–Welzel axes using binary contrastive vignettes.
PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
PhysStream is an autoregressive image-to-video generation method for physics-grounded, interactive control in multi-object tabletop scenes.
World Models for Embodied Intelligence: From Plausible to Controllable to Actionable
This paper is a survey of world models for embodied intelligence that argues they should be judged by how they improve behavior rather than by visual realism alone.
Archive
Weekly Archive
15自己回帰型動画のためのメモリと表現
Work on autoregressive video models centers on two linked problems: preserving useful visual history over long rollouts and learning representations that keep generation efficient and coherent.
オープンLLMのポストトレーニングと評価
The theme centers on open, multi-stage LLM post-training in which evaluation and verification guide data selection, preference tuning, and reinforcement learning.
適応的マルチモーダル表現学習
The theme centers on adapting representations to heterogeneous visual and linguistic inputs rather than relying on uniform processing.
LLM KVキャッシュ圧縮
This theme centers on making long-context LLM inference practical by reducing the rapidly growing KV cache memory footprint.
複雑系のための基盤型ワールドモデル
This theme centers on moving beyond direct answer generation toward models that simulate underlying system dynamics before making decisions or predictions.
全二重対話の評価
This theme centers on evaluating speech-language systems that listen and speak concurrently, moving beyond conventional turn-based assessment.
統合マルチモーダル表現学習
This week's theme centers on how multimodal models should represent visual information when a single system is expected to understand, generate, and edit images.
頑健な音声認識システム
This theme covers speech recognition systems designed to operate beyond clean single-speaker conditions, addressing noisy or soundproof environments, overlapping speakers, and channel mismatch.
リスク制御された自己改善
This theme centers on how learning systems can improve themselves or revise their own reasoning without silently accumulating harmful changes.
翻訳・エージェント向け評価ベンチマーク
This week's theme centers on evaluation shifting toward more realistic, high-constraint benchmarks rather than simplified test settings.
3D再構成と動的ガウシアンモデリング
This theme centers on making 3D/4D reconstruction and rendering more practical under sparse-view or monocular conditions by moving beyond purely photometric fitting.
強化学習における目的関数設計
This theme centers on rethinking what reinforcement learning should optimize, beyond standard expected-return training.
マルチモーダル推論の評価
This theme focuses on evaluating whether language models and multimodal systems can genuinely reason across visual, linguistic, and spatial representations rather than rely on superficial benchmark signals.
オムニモーダル表現と評価
This theme centers on models that unify multiple modalities while making their representations more controllable and more meaningfully testable.
エージェント環境とハーネスのスケーリング
This theme centers on the shift from evaluating agents in fixed setups to scaling the environments, interaction loops, and runtime scaffolds that shape agent behavior.