STRAPSim: A Portfolio Similarity Metric for ETF Alignment and Portfolio Trades
- URL: http://arxiv.org/abs/2509.24151v1
- Date: Mon, 29 Sep 2025 00:57:41 GMT
- Title: STRAPSim: A Portfolio Similarity Metric for ETF Alignment and Portfolio Trades
- Authors: Mingshu Li, Dhruv Desai, Jerinsh Jeyapaulraj, Philip Sommer, Riya Jain, Peter Chu, Dhagash Mehta,
- Abstract summary: STRAPSim is a novel method that computes portfolio similarity by matching constituents based on semantic similarity.<n>We benchmark our approach against Jaccard, weighted Jaccard, as well as BERTScore-inspired variants across public classification, regression, and recommendation tasks.<n> Empirical results show that our method consistently outperforms baselines in predictive accuracy and ranking alignment.
- Score: 0.5847369405576658
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Accurately measuring portfolio similarity is critical for a wide range of financial applications, including Exchange-traded Fund (ETF) recommendation, portfolio trading, and risk alignment. Existing similarity measures often rely on exact asset overlap or static distance metrics, which fail to capture similarities among the constituents (e.g., securities within the portfolio) as well as nuanced relationships between partially overlapping portfolios with heterogeneous weights. We introduce STRAPSim (Semantic, Two-level, Residual-Aware Portfolio Similarity), a novel method that computes portfolio similarity by matching constituents based on semantic similarity, weighting them according to their portfolio share, and aggregating results via residual-aware greedy alignment. We benchmark our approach against Jaccard, weighted Jaccard, as well as BERTScore-inspired variants across public classification, regression, and recommendation tasks, as well as on corporate bond ETF datasets. Empirical results show that our method consistently outperforms baselines in predictive accuracy and ranking alignment, achieving the highest Spearman correlation with return-based similarity. By leveraging constituent-aware matching and dynamic reweighting, portfolio similarity offers a scalable, interpretable framework for comparing structured asset baskets, demonstrating its utility in ETF benchmarking, portfolio construction, and systematic execution.
Related papers
- Cross-Sectional Asset Retrieval via Future-Aligned Soft Contrastive Learning [87.54084417547621]
We argue that effective asset retrieval should be future-aligned.<n>Experiments on 4,229 US equities demonstrate that Future-Aligned Soft Contrastive Learning consistently outperforms 13 baselines across all future-behavior metrics.
arXiv Detail & Related papers (2026-02-11T10:17:52Z) - A Novel approach to portfolio construction [0.0]
This paper proposes a machine learning-based framework for asset selection and portfolio construction.<n>It is called the Best-Path Algorithm Sparse Graphical Model (BPASGM)<n>Monte Carlo simulations show BPASGM-based portfolios achieve more stable risk-return profiles, lower realized volatility, and superior risk-adjusted performance.
arXiv Detail & Related papers (2026-02-03T09:52:06Z) - Diffolio: A Diffusion Model for Multivariate Probabilistic Financial Time-Series Forecasting and Portfolio Construction [7.0782219254786725]
We propose Diffolio, a diffusion model designed for multivariate financial time-series forecasting and portfolio construction.<n>Diffolio employs a denoising network with a hierarchical attention architecture, comprising both asset-level and market-level layers.
arXiv Detail & Related papers (2025-11-10T12:05:32Z) - From Headlines to Holdings: Deep Learning for Smarter Portfolio Decisions [4.288926547930663]
We present an end-to-end framework that learns portfolio weights using deep learning.<n>We evaluate the framework on nine U.S. stocks spanning six sectors, chosen to balance sector diversity and news coverage.<n>Although the stock universe is limited, the results underscore the value of integrating price, relational, and sentiment signals for portfolio management.
arXiv Detail & Related papers (2025-09-29T00:42:24Z) - TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them [58.04324690859212]
Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks.<n>We identify two fundamental types of inconsistencies: Score-Comparison Inconsistency and Pairwise Transitivity Inconsistency.<n>We propose TrustJudge, a probabilistic framework that addresses these limitations through two key innovations.
arXiv Detail & Related papers (2025-09-25T13:04:29Z) - RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging [69.2230254959204]
We propose RouteMark, a framework for IP protection in merged MoE models.<n>Our key insight is that task-specific experts exhibit stable and distinctive routing behaviors under probing inputs.<n>For attribution and tampering detection, we introduce a similarity-based matching algorithm.
arXiv Detail & Related papers (2025-08-03T14:51:58Z) - NDCG-Consistent Softmax Approximation with Accelerated Convergence [67.10365329542365]
We propose novel loss formulations that align directly with ranking metrics.<n>We integrate the proposed RG losses with the highly efficient Alternating Least Squares (ALS) optimization method.<n> Empirical evaluations on real-world datasets demonstrate that our approach achieves comparable or superior ranking performance.
arXiv Detail & Related papers (2025-06-11T06:59:17Z) - Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis [89.60263788590893]
Post-training Quantization (PTQ) technique has been extensively adopted for large language models (LLMs) compression.<n>Existing algorithms focus primarily on performance, overlooking the trade-off among model size, performance, and quantization bitwidth.<n>We provide a novel benchmark for LLMs PTQ in this paper.
arXiv Detail & Related papers (2025-02-18T07:35:35Z) - Supervised Similarity for High-Yield Corporate Bonds with Quantum Cognition Machine Learning [0.8706730566331037]
We investigate the application of quantum cognition machine learning (QCML) to distance metric learning in corporate bond markets.<n>We show that QCML outperforms classical tree-based models in high-yield (HY) markets, while giving comparable or better performance in investment grade (IG) markets.
arXiv Detail & Related papers (2025-02-03T16:28:44Z) - Optimizing Portfolio Performance through Clustering and Sharpe Ratio-Based Optimization: A Comparative Backtesting Approach [0.0]
This paper introduces a comparative backtesting approach that combines clustering-based portfolio segmentation and Sharpe ratio-based optimization to enhance investment decision-making.<n>We segment a diverse set of financial assets into clusters based on their historical log-returns using K-Means clustering.<n>For each cluster, we apply a Sharpe ratio-based optimization model to derive optimal weights that maximize risk-adjusted returns.
arXiv Detail & Related papers (2025-01-21T12:00:52Z) - Interpretable Company Similarity with Sparse Autoencoders [0.0]
Sparse Autoencoders (SAEs) have shown promise in enhancing the interpretability of Large Language Models (LLMs)<n>We benchmark SAE features against SIC-codes, Industry codes, and Embeddings.<n>Our results demonstrate SAE features surpass sector classifications and embeddings in capturing fundamental company characteristics.
arXiv Detail & Related papers (2024-12-03T17:34:50Z) - Top-K Pairwise Ranking: Bridging the Gap Among Ranking-Based Measures for Multi-Label Classification [120.37051160567277]
This paper proposes a novel measure named Top-K Pairwise Ranking (TKPR)
A series of analyses show that TKPR is compatible with existing ranking-based measures.
On the other hand, we establish a sharp generalization bound for the proposed framework based on a novel technique named data-dependent contraction.
arXiv Detail & Related papers (2024-07-09T09:36:37Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.