論文の概要: Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference
- arxiv url: http://arxiv.org/abs/2610.00586v1
- Date: Wed, 30 Sep 2026 18:51:30 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.721745
- Title: Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference
- Title(参考訳): Right In-Place (RiP) Convolution: メモリ効率の良いCNN推論のためのシンプルで汎用的で、ほぼ最適戦略
- Abstract要約: 演算ではなくアクティベーションメモリは、マイクロコントローラのような制約のあるハードウェア上でのCNN推論を制限する。
直接 in-place 畳み込みは二重バッファのコストを低減させるが、Gural と Murmann のメモリ最適化の定式化は有効なパディング、単位非有界ストライド、単位拡張、奇数正方核を仮定する。
本稿では,各レイヤが共有ワークスペース内で入力を右に読み取り,インデックスゼロから出力を左に書き込むという,ビット単位の演算であるRight In-Place畳み込みを提案する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Activation memory, not compute, limits CNN inference on constrained hardware such as microcontrollers. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann assumes valid padding, unit stride, unit dilation, and odd square kernels, and needs a non-sequential traversal costing $2\times$ inference time in transposes. We identify two regimes in which their published closed form does not hold: (1) an under-allocation of exactly $(k-1)C_{in} \bmod (C_{out}-C_{in})$ scalars, active on every convolutional layer of their own deployed network and manifesting as a silent corruption of still-live input; (2) an unbounded overestimate, up to $2{,}432\times$, once the critical leg leaves the output grid. We correct both and generalize to arbitrary stride, dilation, padding, and rectangular kernels. We then propose Right In-Place (RiP) convolution, a bit-identical operation in which every layer reads its input right-aligned in a shared workspace and writes its output left-aligned from index zero. The debt is piecewise affine in the output pixel index, so evaluating its breakpoints in $O(1)$ yields the minimum safe gap without enumerating the output grid, with row-major access preserved. Across $10{,}000$ random layers RiP produced no corruption, and across 84 convolutional layers from 25 architectures it matches the herringbone workspace exactly on 58 and within 5% on 81, using 24.8% less memory than dual buffering on average. Written into TinyEngine's kernels and deployed to a Raspberry Pi Pico 1 and Pico 2, it cuts peak activation memory across eleven MCUNet models by 12.5 to 33.3% at unchanged cycle counts and bit-identical outputs, raising the number of models that fit the Pico 1's 256 KB SRAM from six to nine.
- Abstract(参考訳): 演算ではなくアクティベーションメモリは、マイクロコントローラのような制約のあるハードウェア上でのCNN推論を制限する。
直接 in-place 畳み込みは二重バッファのコストを除去するが、Gural と Murmann のメモリ最適化の定式化は、有効なパディング、ユニットストライド、ユニットディレーション、奇数正方核を仮定し、非逐次的トラバースのコストは2\times$ inference time in transposes である。
1) 正確な$(k-1)C_{in} \bmod (C_{out}-C_{in})$ scalars, active on all convolutional layer of their deployed network and manifesting as a silent corruption of still-live input; (2) unbounded overetimate, up to $2{,}432\times$。
我々はどちらも修正し、任意のストライド、拡張、パディング、矩形カーネルに一般化する。
次に、Right In-Place(RiP)畳み込み(Right In-Place)を提案する。
借金は出力ピクセルインデックスに一括アフィンなので、そのブレークポイントを$O(1)$で評価すると、出力グリッドを列挙することなく最小の安全なギャップが得られ、行長アクセスが保存される。
10{,}000$のランダムなレイヤで、25のアーキテクチャから84の畳み込みレイヤにまたがって、平均して2つのバッファリングよりも24.8%少ないメモリを使用して、58と81の5%のハーリングボーンワークスペースと正確に一致した。
TinyEngineのカーネルに書き込まれ、Raspberry Pi Pico 1とPico 2にデプロイされ、11のMCUNetモデルでピークアクティベーションメモリを12.5から33.3%削減した。
関連論文リスト
- BiKAN: Restoring Collapsed Basis of Binary Kolmogorov--Arnold Networks [6.860988566886594]
Kolmogorov-Arnold Network(KAN)のバイナリ化は各レイヤで利用可能な関数空間を変更する。
提案するBiKANは,各2進数2のWalsh文字を加算することにより,この問題に対処する。
W1A1では、ビカンは99.48%、84.38%、Mは55.81%、401AR-10は72AR-100である。
論文 参考訳(メタデータ) (2026-08-02T20:49:59Z) - Scaling Neural Network Verification with Tensor Parallelism and Fully Sharded Data Parallelism [0.0]
形式的ニューラルネットワーク検証は、実際にはGPUメモリによって境界付けられている。
大規模なモデルトレーニングのために開発された2つのテクニックをauto_LiRPA / $,$-CROWN 検証フレームワークに適用する。
フルシャードデータ並列(FSDP)シャードは、層ごとのAllGatherで重量行列のみをシャードし、単一GPUベースラインとビット単位で同一のバウンドを生成する。
論文 参考訳(メタデータ) (2026-06-08T11:56:29Z) - SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data [3.1624024957575982]
SURGEは,4万の論理パーティションに8億以上のテキストの埋め込みを生成するために,本番環境にデプロイされたストリーミングエンコーディングシステムである。
4つのNVIDIA L4768を持つ10Mテキストでは、SURGEは26,413のテキスト/sを提供する。
論文 参考訳(メタデータ) (2026-05-01T19:51:50Z) - Scaling DoRA: High-Rank Adaptation via Factored Norms and Fused Kernels [83.99688944263843]
DoRA(Weight-De Low-Rank Adaptation)は、LoRAを方向から分離することで拡張する。
d_in = 8192 とランク r = 384 では、単一のモジュールのノルムは bf16 で512MB の過渡的なワーキングメモリを必要とする。
因子ノルムは、二乗ノルムを O(d_out r + r2) 中間体を通して計算可能な基底、交差、およびグラマー項に分解し、密積を除去する。
論文 参考訳(メタデータ) (2026-03-23T17:57:24Z) - FHECore: Rethinking GPU Microarchitecture for Fully Homomorphic Encryption [2.7777199166440827]
FHE(Fully Homomorphic Encryption)は、暗号化されたデータを直接計算できるが、膨大な計算とメモリオーバーヘッドを発生させる。
カスタムアクセラレーターはこれらのコストを軽減することができるが、市場投入までの長い時間とFHEアルゴリズムの急速な進化は、長期的な妥当性を脅かす。
本稿では,GPUのストリームマルチプロセッサに直接統合された特殊な機能ユニットであるFHECoreを提案する。
論文 参考訳(メタデータ) (2026-02-10T02:55:10Z) - Cut Your Losses in Large-Vocabulary Language Models [102.6981011879656]
我々は,全トークンのロジットをグローバルメモリに実体化することなく,クロスエントロピー損失を計算する手法であるカットクロスエントロピー(CCE)を提案する。
CCEはロスのメモリフットプリントを24GBから1MBに減らし、ヘッドのトレーニング時間のメモリ消費を28GBから1GBに短縮する。
論文 参考訳(メタデータ) (2024-11-13T20:30:15Z) - Scalable 3D Registration via Truncated Entry-wise Absolute Residuals [65.04922801371363]
3ドルの登録アプローチでは、1000万ドル(107ドル)以上のポイントペアを、99%以上のランダムなアウトレイアで処理することができる。
我々はこの手法をTEARと呼び、Trncated Entry-wise Absolute Residualsを演算するoutlier-robust損失を最小限にする。
論文 参考訳(メタデータ) (2024-04-01T04:43:39Z) - MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning [72.80896338009579]
メモリボトルネックは畳み込みニューラルネットワーク(CNN)の設計における不均衡なメモリ分布に起因する。
本稿では,ピークメモリを大幅に削減するパッチ・バイ・パッチ・推論スケジューリングを提案する。
ニューラルアーキテクチャサーチによるプロセスを自動化し、ニューラルアーキテクチャと推論スケジューリングを共同で最適化し、MCUNetV2に導いた。
論文 参考訳(メタデータ) (2021-10-28T17:58:45Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。