論文の概要: S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing
- arxiv url: http://arxiv.org/abs/2606.24441v1
- Date: Tue, 23 Jun 2026 11:19:34 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-24 22:16:48.916028
- Title: S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing
- Title(参考訳): S1-Omni画像:科学画像理解・生成・編集のための統一モデル
- Authors: Qingxiao Li, Zikai Wang, Qingli Wang, Nan Xu,
- Abstract要約: 本稿では,科学画像の理解,生成,編集のためのオープンウェイト統一マルチモーダルモデルを提案する。
S1-Omni-Imageは、バックボーンS1-VL-32Bの科学的なマルチモーダル推論に基づいている。
統一されたフレームワークで、科学的な画像理解、生成、編集をサポートする。
- 参考スコア(独自算出の注目度): 25.725824689793402
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose image generation models, scientific image tasks require not only high-fidelity synthesis, but also robust understanding of scientific semantics, structural relations, domain knowledge, and task intent. To this end, S1-Omni-Image builds on the scientific multimodal reasoning backbone S1-VL-32B and couples its understanding capability with an image generation module under a unified think-before-generate paradigm. Given a user instruction, the model first produces a task-oriented reasoning trace, a textual answer, and a task special token; their hidden states are then injected into the generation module to condition image generation or editing. S1-Omni-Image supports scientific image understanding, generation, and editing in a unified framework. For generation, it focuses on scientific illustrations and text rendering, including logical diagrams, relational comparisons, data charts, and realistic scientific visualizations. For editing, it casts segmentation and other domain-specific vision tasks as native image editing problems, enabling multi-turn illustration editing, medical and geographic image segmentation, medical image translation, and scientific image super-resolution. We construct SciGenEdit, a 314K-sample training dataset, and release the model weights, inference code, and SciGenEdit-10K. Experiments show that S1-Omni-Image substantially improves scientific image generation and editing while preserving the scientific image understanding capability inherited from S1-VL-32B. It outperforms open-source models on GenExam and TechImage-Bench, achieves state-of-the-art results on four editing benchmarks including MSD, cigRockSEM, SynthRAD2025, and IXI, and maintains stable performance on scientific image understanding evaluations.
- Abstract(参考訳): S1-Omni-Imageは、科学画像の理解、生成、編集のためのオープンウェイト統一マルチモーダルモデルである。
汎用画像生成モデルとは異なり、科学的画像タスクは高忠実性合成だけでなく、科学的意味論、構造的関係、ドメイン知識、タスク意図の堅牢な理解も必要である。
この目的のために、S1-Omni-Imageは科学的なマルチモーダル推論バックボーンS1-VL-32Bに基づいて構築され、その理解能力と画像生成モジュールを統合思考前生成パラダイムの下で結合する。
ユーザ命令が与えられたら、まずタスク指向の推論トレース、テキスト応答、タスク特殊トークンを生成し、その隠れた状態を生成モジュールに注入して画像生成や編集を行う。
S1-Omni-Imageは、統一されたフレームワークで科学的な画像理解、生成、編集をサポートする。
世代別では、論理図、関係比較、データチャート、現実的な科学的視覚化など、科学的イラストやテキストレンダリングに重点を置いている。
編集には、セグメンテーションやその他のドメイン固有の視覚タスクをネイティブな画像編集問題として採用し、マルチターンイラストレーション編集、医療画像のセグメンテーション、医用画像翻訳、科学画像の超解像を可能にする。
314KサンプルのトレーニングデータセットであるSciGenEditを構築し、モデル重み、推論コード、SciGenEdit-10Kをリリースする。
実験により、S1-Omni-ImageはS1-VL-32Bから受け継いだ科学的画像理解能力を保ちながら、科学的画像の生成と編集を大幅に改善することが示された。
GenExamとTechImage-Benchのオープンソースモデルより優れており、MSD、cigRockSEM、SynthRAD2025、IXIの4つの編集ベンチマークで最先端の結果が得られ、科学画像理解評価における安定したパフォーマンスを維持している。
関連論文リスト
- Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models [49.93401412148884]
FEPBenchは、慎重に選択された高品質な科学イラストから構築されたベンチマークである。
我々は,T2Iモデルについて,命令忠実度,推論エンリッチメント,意味的精度の3つの次元に沿って評価する。
その結果、GPT Image 2やNano Banana Proのような最先端のクローズドソースモデルでさえ、テキストレンダリングのボトルネック、推論のリッチ化の制限、生成のリッチネスと精度のバランスの難しさに悩まされていることがわかった。
論文 参考訳(メタデータ) (2026-06-04T09:49:02Z) - HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer [104.09730595701468]
画素空間拡散変換器を用いた統合生成基盤モデルであるHiDream-O1-Imageを提案する。
HiDream-O1-Imageは、原画像ピクセル、テキストトークン、タスク固有の条件を単一の共有トークン空間にマッピングすることにより、マルチモーダル入力の構造的統一を実現する。
実験により、HiDream-O1-Imageは、テキスト・ツー・イメージ生成、命令ベースの編集、主観的パーソナライゼーションなど、さまざまな世代のタスクに優れることが示された。
論文 参考訳(メタデータ) (2026-05-11T17:59:09Z) - SIQA: Toward Reliable Scientific Image Quality Assessment [72.41803245808924]
我々は,2つの相補的な次元に沿って,科学的画質をモデル化するフレームワークであるSIQA(Scientific Image Quality Assessment)を紹介する。
SIQA-U (Understanding), SIQA-S (Scoring), SIQA-U (Understanding), SIQA-U (Understanding), SIQA-U (Understanding), SIQA-U (Understanding), SIQA-U (Understanding) の2つの評価プロトコルを設計した。
代表的マルチモーダル大言語モデル(MLLM)に対する実験は、アライメントアライメントと科学的理解の間に一貫した相違が見られる。
論文 参考訳(メタデータ) (2026-03-05T06:57:26Z) - UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing [44.071171929398076]
マルチモーダルモデルは、しばしば深い推論を必要とする複雑な合成タスクに苦しむ。
画像生成と画像編集を調和させる統一フレームワークUniReasonを提案する。
我々は,大規模推論中心のデータセットを体系的に構築することで,このフレームワークをサポートする。
論文 参考訳(メタデータ) (2026-02-02T18:34:35Z) - Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility [57.83550091882176]
生成パラダイム,評価,下流利用における科学的画像合成について検討する。
本稿では,情報の有用性と論理的妥当性に基づいて生成した画像を評価するSciGenBenchを紹介する。
厳密に検証された合成科学画像上の微調整された大規模マルチモーダルモデルにより、一貫した推論ゲインが得られることを示す。
論文 参考訳(メタデータ) (2026-01-17T14:18:36Z) - S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding [16.351123624587384]
S1-MMAlignは1550万以上の高品質の画像テキストペアからなる大規模で多分野のマルチモーダルデータセットである。
本稿では,Qwen-VL多モード大モデル系列を用いたAI対応セマンティックエンハンスメントパイプラインを提案する。
論文 参考訳(メタデータ) (2026-01-01T08:54:51Z) - Charts Are Not Images: On the Challenges of Scientific Chart Editing [66.38730113476677]
textitFigEditは、3万以上のサンプルからなる科学的フィギュア編集のベンチマークである。
私たちのベンチマークでは、ピクセルレベルの操作の重大な制限が示されています。
textitFigEdit をリリースすることにより,構造対応図形編集の体系的な進歩の実現を目指す。
論文 参考訳(メタデータ) (2025-11-30T06:13:48Z) - SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model [21.81341169834812]
SridBenchは、科学フィギュア生成のための最初のベンチマークである。
これは13の自然科学とコンピュータ科学の分野にわたる主要な科学論文から1,120の事例で構成されている。
その結果、GPT-4o画像のような最上位モデルでさえ、人間のパフォーマンスに遅れがあることが判明した。
論文 参考訳(メタデータ) (2025-05-28T08:51:01Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。