The Silent Spill: Measuring Sensitive Data Leaks Across Public URL Repositories
- URL: http://arxiv.org/abs/2602.21826v1
- Date: Wed, 25 Feb 2026 11:54:46 GMT
- Title: The Silent Spill: Measuring Sensitive Data Leaks Across Public URL Repositories
- Authors: Tarek Ramadan, AbdelRahman Abdou, Mohammad Mannan, Amr Youssef,
- Abstract summary: We present an automated system that detects and analyzes potential sensitive information leaked through publicly accessible URLs.<n>We apply it to 6,094,475 URLs collected from public scanning platforms, paste sites, and web archives.<n>These findings show that sensitive information remains exposed, underscoring the importance of automated detection to identify accidental leaks.
- Score: 3.557034943202842
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: A large number of URLs are made public by various platforms for security analysis, archiving, and paste sharing -- such as VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt. These services may unintentionally expose links containing sensitive information, as reported in some news articles and blog posts. However, no large-scale measurement has quantified the extent of such exposures. We present an automated system that detects and analyzes potential sensitive information leaked through publicly accessible URLs. The system combines lexical URL filtering, dynamic rendering, OCR-based extraction, and content classification to identify potential leaks. We apply it to 6,094,475 URLs collected from public scanning platforms, paste sites, and web archives, identifying 12,331 potential exposures across authentication, financial, personal, and document-related domains. These findings show that sensitive information remains exposed, underscoring the importance of automated detection to identify accidental leaks.
Related papers
- Analyzing the Availability of E-Mail Addresses for PyPI Libraries [89.21869606965578]
81.6% of libraries include at least one valid e-mail address, with PyPI serving as the primary source.<n>We identify over 698,000 invalid entries, primarily due to missing fields.
arXiv Detail & Related papers (2026-01-20T14:54:58Z) - Characterizing Phishing Pages by JavaScript Capabilities [77.64740286751834]
This paper aims to aid researchers and analysts by automatically differentiating groups of phishing pages based on the underlying kit.<n>For kit detection, our system has an accuracy of 97% on a ground-truth dataset of 548 kit families deployed across 4,562 phishing URLs.<n>We find that UI interactivity and basic fingerprinting are universal techniques, present in 90% and 80% of the clusters.
arXiv Detail & Related papers (2025-09-16T15:39:23Z) - LLM-Based Identification of Infostealer Infection Vectors from Screenshots: The Case of Aurora [0.0]
Infostealers exfiltrate credentials, session cookies, and sensitive data from infected systems.<n>With over 29 million stealer logs reported in 2024, manual analysis and mitigation at scale are virtually unfeasible/unpractical.<n>This paper introduces a novel approach leveraging Large Language Models (LLMs) to analyze infection screenshots.
arXiv Detail & Related papers (2025-07-31T14:49:03Z) - Client-Side Zero-Shot LLM Inference for Comprehensive In-Browser URL Analysis [0.0]
Malicious websites and phishing URLs pose an ever-increasing cybersecurity risk.<n>Traditional detection approaches rely on machine learning.<n>We propose a novel client-side framework for comprehensive URL analysis.
arXiv Detail & Related papers (2025-06-04T07:47:23Z) - Automated Profile Inference with Language Model Agents [67.32226960040514]
We study a new threat that LLMs pose to online pseudonymity, called automated profile inference.<n>An adversary can instruct LLMs to automatically scrape and extract sensitive personal attributes from publicly visible user activities on pseudonymous platforms.<n>We introduce an automated profiling framework called AutoProfiler to assess the feasibility of such threats in real-world scenarios.
arXiv Detail & Related papers (2025-05-18T13:05:17Z) - Document Screenshot Retrievers are Vulnerable to Pixel Poisoning Attacks [72.4498910775871]
Vision-language model (VLM)-based retrievers leverage document screenshots embedded as vectors to enable effective search and offer a simplified pipeline over traditional text-only methods.<n>In this study, we propose three pixel poisoning attack methods designed to compromise VLM-based retrievers.
arXiv Detail & Related papers (2025-01-28T12:40:37Z) - SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model [48.547599530927926]
Synthetic images, when shared on social media, can mislead extensive audiences and erode trust in digital content.<n>We introduce the Social media Image Detection dataSet (SID-Set), which offers three key advantages.<n>We propose a new image deepfake detection, localization, and explanation framework, named SIDA.
arXiv Detail & Related papers (2024-12-05T16:12:25Z) - Can Features for Phishing URL Detection Be Trusted Across Diverse Datasets? A Case Study with Explainable AI [0.0]
Phishing has been a prevalent cyber threat that manipulates users into revealing sensitive private information through deceptive tactics.
proactively detection of phishing URLs (or websites) has been established as an widely-accepted defense approach.
We analyze two publicly available phishing URL datasets, where each dataset has its own set of unique and overlapping features related to URL string and website contents.
arXiv Detail & Related papers (2024-11-14T21:07:52Z) - Automatic Generation of Web Censorship Probe Lists [6.051603326423421]
Previous efforts to generate domain probe lists have been mostly manual or crowdsourced.
This paper explores methods for automatically generating probe lists that are both comprehensive and up-to-date for Web censorship measurement.
arXiv Detail & Related papers (2024-07-11T05:04:52Z) - An Adversarial Attack Analysis on Malicious Advertisement URL Detection
Framework [22.259444589459513]
Malicious advertisement URLs pose a security risk since they are the source of cyber-attacks.
Existing malicious URL detection techniques are limited and to handle unseen features as well as generalize to test data.
In this study, we extract a novel set of lexical and web-scrapped features and employ machine learning technique to set up system for fraudulent advertisement URLs detection.
arXiv Detail & Related papers (2022-04-27T20:06:22Z) - Exposing Query Identification for Search Transparency [69.06545074617685]
We explore the feasibility of approximate exposing query identification (EQI) as a retrieval task by reversing the role of queries and documents in two classes of search systems.
We derive an evaluation metric to measure the quality of a ranking of exposing queries, as well as conducting an empirical analysis focusing on various practical aspects of approximate EQI.
arXiv Detail & Related papers (2021-10-14T20:19:27Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.