論文の概要: Leveraging Large Language Models for Trustworthiness Assessment of Web Applications
- arxiv url: http://arxiv.org/abs/2603.23781v1
- Date: Tue, 24 Mar 2026 23:33:54 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-26 21:06:11.053769
- Title: Leveraging Large Language Models for Trustworthiness Assessment of Web Applications
- Title(参考訳): Webアプリケーションの信頼性評価のための大規模言語モデルの活用
- Authors: Oleksandr Yarotskyi, José D'Abruzzo Pereira, João R. Campos,
- Abstract要約: 本研究では,大規模言語モデル(LLM)を活用したWebアプリケーションの信頼性評価を自動化する実証的手法を提案する。
本稿では,LSP(Logic Score of Preference)に基づく階層品質モデルの拡張を提案する。
実験結果から,過度な構造的コンテキストがノイズを発生させる可能性が示唆された。
- 参考スコア(独自算出の注目度): 13.909850314037653
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: The widespread adoption of web applications has made their security a critical concern and has increased the need for systematic ways to assess whether they can be considered trustworthy. However, "trust" assessment remains an open problem as existing techniques primarily focus on detecting known vulnerabilities or depend on manual evaluation, which limits their scalability; therefore, evaluating adherence to secure coding practices offers a complementary, pragmatic perspective by focusing on observable development behaviors. In practice, the identification and verification of secure coding practices are predominantly performed manually, relying on expert knowledge and code reviews, which is time-consuming, subjective, and difficult to scale. This study presents an empirical methodology to automate the trustworthiness assessment of web applications by leveraging Large Language Models (LLMs) to verify adherence to secure coding practices. We conduct a comparative analysis of prompt engineering techniques across five state-of-the-art LLMs, ranging from baseline zero-shot classification to prompts enriched with semantic definitions, structural context derived from call graphs, and explicit instructional guidance. Furthermore, we propose an extension of a hierarchical Quality Model (QM) based on the Logic Score of Preference (LSP), in which LLM outputs are used to populate the model's quality attributes and compute a holistic trustworthiness score. Experimental results indicate that excessive structural context can introduce noise, whereas rule-based instructional prompting improves assessment reliability. The resulting trustworthiness score allows discriminating between secure and vulnerable implementations, supporting the feasibility of using LLMs for scalable and context-aware trust assessment.
- Abstract(参考訳): Webアプリケーションの普及により、セキュリティが重要な問題となり、信頼に値するかどうかを評価するための体系的な方法の必要性が高まっている。
しかし、既存の技術は既知の脆弱性の検出や、そのスケーラビリティを制限する手作業による評価に重点を置いているため、"信頼"評価は未解決の問題のままである。
実際に、セキュアなコーディングプラクティスの識別と検証は、主に手作業で行われ、専門家の知識とコードレビューに依存します。
本研究では,Large Language Models (LLMs) を利用して,Webアプリケーションの信頼性評価を自動化する実験手法を提案する。
基礎となるゼロショット分類から意味定義に富んだプロンプト,コールグラフから派生した構造的コンテキスト,明示的な指導指導まで,5つの最先端LLMにおけるプロンプトエンジニアリング技術の比較分析を行う。
さらに,LSP(Logic Score of Preference)に基づく階層的品質モデル(QM)の拡張を提案する。
実験結果から,過度な構造的コンテキストがノイズを発生させる可能性が示唆された。
結果として生じる信頼性スコアは、セキュアな実装と脆弱な実装の区別を可能にし、スケーラブルでコンテキスト対応の信頼評価にLLMを使用することの可能性をサポートする。
関連論文リスト
- OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics [82.0813150432867]
我々は,大規模言語モデル(LLM)のアンラーニング手法とメトリクスをベンチマークするための標準フレームワークであるOpenUnlearningを紹介する。
OpenUnlearningは、13のアンラーニングアルゴリズムと16のさまざまな評価を3つの主要なベンチマークで統合する。
また、多様なアンラーニング手法をベンチマークし、広範囲な評価スイートとの比較分析を行う。
論文 参考訳(メタデータ) (2025-06-14T20:16:37Z) - SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis [39.229080120880774]
SV-TrustEval-Cは,C言語で記述されたコードの脆弱性解析のための大規模言語モデルの能力を評価するためのベンチマークである。
以上の結果から,現在のLLMは複雑なコード関係を理解するのに十分ではないことが示され,その脆弱性分析はロバストな論理的推論よりもパターンマッチングに頼っている。
論文 参考訳(メタデータ) (2025-05-27T02:16:27Z) - Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling [41.19330514054401]
大規模言語モデル(LLM)は、不一致の自己認識に起因する幻覚の傾向にある。
本稿では,高速かつ低速な推論システムを統合し,信頼性とユーザビリティを調和させる明示的知識境界モデリングフレームワークを提案する。
論文 参考訳(メタデータ) (2025-03-04T03:16:02Z) - The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? [1.3810901729134184]
大きな言語モデル(LLM)は、真の言語理解と適応性を示すのに失敗しながら、標準化されたテストで優れている。
NLP評価フレームワークの系統的解析により,評価スペクトルにまたがる広範囲にわたる脆弱性が明らかになった。
我々は、操作に抵抗し、データの汚染を最小限に抑え、ドメイン固有のタスクを評価する新しい評価方法の土台を築いた。
論文 参考訳(メタデータ) (2024-12-02T20:49:21Z) - SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts [0.6291443816903801]
本稿では,大規模言語モデル(LLM)のロバスト性を自律的に評価する新しいフレームワークを提案する。
本稿では,ドメイン制約付き知識グラフ三重項から記述文を生成し,敵対的プロンプトを定式化する。
この自己評価機構により、LCMは外部ベンチマークを必要とせずにその堅牢性を評価することができる。
論文 参考訳(メタデータ) (2024-12-01T10:58:53Z) - Unveiling the Misuse Potential of Base Large Language Models via In-Context Learning [61.2224355547598]
大規模言語モデル(LLM)のオープンソース化は、アプリケーション開発、イノベーション、科学的進歩を加速させる。
我々の調査は、この信念に対する重大な監視を露呈している。
我々の研究は、慎重に設計されたデモを配置することにより、ベースLSMが悪意のある命令を効果的に解釈し実行できることを実証する。
論文 参考訳(メタデータ) (2024-04-16T13:22:54Z) - TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness [58.721012475577716]
大規模言語モデル(LLM)は、様々な領域にまたがる印象的な能力を示しており、その実践的応用が急増している。
本稿では,行動整合性の概念に基づくフレームワークであるTrustScoreを紹介する。
論文 参考訳(メタデータ) (2024-02-19T21:12:14Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。