論文の概要: An Extensive Empirical Study on Evaluation Metrics for Combinatorial Interaction Testing
- arxiv url: http://arxiv.org/abs/2610.02560v1
- Date: Thu, 01 Oct 2026 22:53:23 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-06 00:14:30.115331
- Title: An Extensive Empirical Study on Evaluation Metrics for Combinatorial Interaction Testing
- Title(参考訳): 組合せインタラクションテストのための評価指標に関する総合的実証的研究
- Abstract要約: Combinatorial Interaction Testing (CIT)は、近年研究と実践の両方で広く注目を集めているブラックボックステスト手法である。
CITテストプロセスの基本的なコンポーネントとして、評価基準はテストスイートの評価と比較において重要な役割を果たす。
適切な測定基準の選択は、これらの特性が評価の有効性に大きな影響を及ぼす可能性があるため、テストスイート固有の特性も考慮すべきである。
- 参考スコア(独自算出の注目度): 10.612267723496851
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Combinatorial interaction testing (CIT) is a black-box testing method that has received extensive attention in both research and practice over recent years. Its primary objective is to construct an effective combinatorial test suite that detects software failures caused by parameter interactions. As a fundamental component of the CIT testing process, the evaluation metric plays a critical role in assessing and comparing combinatorial test suites, as well as in evaluating various test generation techniques. For CIT practitioners, selecting an appropriate evaluation metric is both important and challenging, given the wide variety of available options. Nevertheless, no prior work has systematically addressed this problem. To fill this gap, this paper first provides a comprehensive survey of black-box evaluation metrics for combinatorial test suites, offering rigorous definitions, clear classifications, illustrative examples, and complexity analyses. We then conduct an extensive empirical study involving eight open-source projects, encompassing 32 test scenarios and 295,624 combinatorial test suites. In this study, we examine the correlation between each static evaluation metric and fault-detection effectiveness using two correlation measures. Experimental results show that the Value Combination Coverage (VCC) metric serves as a valid predictor for test-suite evaluation. However, distribution-based metrics generally incur lower computational costs than interaction coverage-based ones. The choice of an appropriate metric should also account for the test suite's inherent properties, as these characteristics can substantially influence the effectiveness of the evaluation. Finally, we provide practical guidelines to assist CIT practitioners in selecting suitable evaluation metrics for assessing or comparing combinatorial test suites.
- Abstract(参考訳): Combinatorial Interaction Testing (CIT)は、近年研究と実践の両方で広く注目を集めているブラックボックステスト手法である。
その主な目的は、パラメータの相互作用に起因するソフトウェア障害を検出する効果的な組合せテストスイートを構築することである。
CITテストプロセスの基本的な構成要素として、評価基準は、組合せテストスイートの評価と比較、および様々なテスト生成技術の評価において重要な役割を果たす。
CIT実践者にとって、さまざまな選択肢を考えると、適切な評価基準を選択することは重要かつ困難である。
それにもかかわらず、この問題を体系的に解決する以前の研究は行われていない。
このギャップを埋めるために,本稿ではまず,厳密な定義,明確な分類,説明例,複雑度解析など,組み合わせテストスイートのブラックボックス評価指標に関する総合的な調査を行う。
次に、32のテストシナリオと295,624の組合せテストスイートを含む8つのオープンソースプロジェクトに関する広範な実証研究を行った。
本研究では,2つの相関測度を用いて,各静的評価基準と断層検出の有効性の相関について検討した。
実験の結果,VCC(Value Combination Coverage)測定値がテストスイート評価の有効な予測因子であることがわかった。
しかし、分布ベースのメトリクスは一般的に、相互作用カバレッジベースのメトリクスよりも計算コストが低い。
適切な測定基準の選択は、これらの特性が評価の有効性に大きな影響を及ぼす可能性があるため、テストスイート固有の特性も考慮すべきである。
最後に,CIT実践者が組合せテストスイートの評価や比較に適する評価指標を選択するのを支援するための実践的ガイドラインを提供する。
関連論文リスト
- TestAgent: An Adaptive and Intelligent Expert for Human Assessment [62.060118490577366]
対話型エンゲージメントによる適応テストを強化するために,大規模言語モデル(LLM)を利用したエージェントであるTestAgentを提案する。
TestAgentは、パーソナライズされた質問の選択をサポートし、テストテイカーの応答と異常をキャプチャし、動的で対話的なインタラクションを通じて正確な結果を提供する。
論文 参考訳(メタデータ) (2025-06-03T16:07:54Z) - NLP and Education: using semantic similarity to evaluate filled gaps in a large-scale Cloze test in the classroom [0.0]
ブラジルの学生を対象にしたクローゼテストのデータを用いて,ブラジルポルトガル語(PT-BR)のWEモデルを用いて意味的類似度を測定した。
WEモデルのスコアと審査員の評価を比較した結果,GloVeが最も効果的なモデルであることが判明した。
論文 参考訳(メタデータ) (2024-11-02T15:22:26Z) - Improving Bias Correction Standards by Quantifying its Effects on Treatment Outcomes [54.18828236350544]
Propensity score matching (PSM) は、分析のために同等の人口を選択することで選択バイアスに対処する。
異なるマッチング手法は、すべての検証基準を満たす場合でも、同じタスクに対する平均処理効果(ATE)を著しく異なるものにすることができる。
この問題に対処するため,新しい指標A2Aを導入し,有効試合数を削減した。
論文 参考訳(メタデータ) (2024-07-20T12:42:24Z) - InterEvo-TR: Interactive Evolutionary Test Generation With Readability
Assessment [1.6874375111244329]
テスタによるインタラクティブな可読性評価をEvoSuiteに組み込むことを提案する。
提案手法であるInterEvo-TRは,検索中に異なるタイミングでテスターと対話する。
その結果,中間結果の選択・提示戦略は可読性評価に有効であることが示唆された。
論文 参考訳(メタデータ) (2024-01-13T13:14:29Z) - Towards Reliable AI: Adequacy Metrics for Ensuring the Quality of
System-level Testing of Autonomous Vehicles [5.634825161148484]
我々は、"Test suite Instance Space Adequacy"(TISA)メトリクスと呼ばれる一連のブラックボックステストの精度指標を紹介します。
TISAメトリクスは、テストスイートの多様性とカバレッジと、テスト中に検出されたバグの範囲の両方を評価する手段を提供する。
AVのシステムレベルのシミュレーションテストにおいて検出されたバグ数との相関を検証し,TISA測定の有効性を評価する。
論文 参考訳(メタデータ) (2023-11-14T10:16:05Z) - Testing the Consistency of Performance Scores Reported for Binary
Classification Problems [0.0]
報告された性能スコアの整合性を評価する数値的手法と推定された実験装置を紹介する。
本研究では,提案手法が不整合を効果的に検出し,研究分野の整合性を保護する方法を示す。
科学コミュニティの利益を得るために、一貫性テストはオープンソースのPythonパッケージで利用可能にしました。
論文 参考訳(メタデータ) (2023-10-19T07:04:29Z) - Position: AI Evaluation Should Learn from How We Test Humans [65.36614996495983]
人間の評価のための20世紀起源の理論である心理測定は、今日のAI評価における課題に対する強力な解決策になり得る、と我々は主張する。
論文 参考訳(メタデータ) (2023-06-18T09:54:33Z) - Better than Average: Paired Evaluation of NLP Systems [31.311553903738798]
評価スコアのインスタンスレベルのペアリングを考慮に入れることの重要性を示す。
平均, 中央値, BT と 2 種類のBT (Elo と TrueSkill) を用いて評価スコアの完全な解析を行うための実用的なツールをリリースする。
論文 参考訳(メタデータ) (2021-10-20T19:40:31Z) - Performance Evaluation of Adversarial Attacks: Discrepancies and
Solutions [51.8695223602729]
機械学習モデルの堅牢性に挑戦するために、敵対攻撃方法が開発されました。
本稿では,Piece-wise Sampling Curving(PSC)ツールキットを提案する。
psc toolkitは計算コストと評価効率のバランスをとるオプションを提供する。
論文 参考訳(メタデータ) (2021-04-22T14:36:51Z) - A Statistical Analysis of Summarization Evaluation Metrics using
Resampling Methods [60.04142561088524]
信頼区間は比較的広く,信頼性の高い自動測定値の信頼性に高い不確実性を示す。
多くのメトリクスはROUGEよりも統計的改善を示していないが、QAEvalとBERTScoreという2つの最近の研究は、いくつかの評価設定で行われている。
論文 参考訳(メタデータ) (2021-03-31T18:28:14Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。