| 研究生: |
簡禎 Chien, Jen |
|---|---|
| 論文名稱: |
以子空間為基礎的商品嵌入概念分離 A subspace-based framework for disentangling concepts in product embeddings |
| 指導教授: | 蕭舜文 |
| 學位類別: |
碩士
Master |
| 系所名稱: |
商學院 - 資訊管理學系 Department of Management Information System |
| 論文出版年: | 2026 |
| 畢業學年度: | 114 |
| 語文別: | 英文 |
| 論文頁數: | 54 |
| 中文關鍵詞: | 商品排序 、候選商品重排序 、概念感知排序 、嵌入概念分離 、概念子空間投影 |
| 外文關鍵詞: | product ranking, candidate re-ranking, concept-aware ranking, embedding disentanglement, concept subspace projection |
| 相關次數: | 點閱:8 下載:0 |
| 分享至: |
| 查詢本校圖書館目錄 查詢臺灣博碩士論文知識加值系統 勘誤回報 |
以稠密嵌入為基礎的排序方法通常使用單一整體向量表示查詢與商品,並能在商品搜 尋中提供良好的語意匹配能力。然而,購物查詢往往同時包含多種概念,例如商品、 目標族群、顏色與材質。當這些概念被壓縮至單一表示中時,特定的查詢條件可能在 相似度計算過程中產生糾纏或被弱化。為解決此限制,本研究提出一個概念感知的商 品排序架構,以多個概念特定相似度訊號補充完整向量相似度。
本研究使用概念詞彙表與主成分分析建構階層式商品、原始商品、目標族群、顏 色與材質等概念子空間,並透過殘差化投影降低不同子空間之間的語意重疊;此外, 建立剩餘語意表示,以保留無法由預先定義概念充分解釋的語意資訊。在候選商品排 序階段,系統依據各查詢中偵測到的屬性啟用相對應的概念特定相似度,並透過最佳 化後的非負權重,將其與原始稠密嵌入相似度進行整合。
在 Amazon Shopping Queries Dataset 上的實驗結果顯示,所提出的方法在不同評估 截點下,相較於稠密嵌入基準皆取得幅度雖小但一致的提升。進一步分析發現,隨著 查詢的語意結構變得更複雜,概念分離所帶來的效益也更加明顯;包含多個同時存在 條件的查詢,其改善幅度高於由單一概念主導的查詢。這些結果顯示,概念特定子空 間能透過強化單一整體嵌入中可能無法充分區分的語意條件,提供互補的排序依據。 因此,本研究所提出的架構是對傳統嵌入式排序方法的補充,而非取代,並提供一種 更具可解釋性、可控制性與可擴充性的概念感知查詢-商品匹配方式。
Dense embedding-based ranking represents a query and a product using a single holistic vec- tor and provides strong semantic matching capabilities for product search. However, shopping queries often contain multiple concepts, such as product, target audience, color, and material. Compressing these concepts into one representation may cause specific query constraints to become entangled or weakened during similarity computation. To address this limitation, this study proposes a concept-aware product ranking framework that complements full-vector sim- ilarity with multiple concept-specific similarity signals.
Concept lexicons and principal component analysis are used to construct subspaces for hi- erarchical product, original product, target audience, color, and material information. Residual- ized projection is applied to reduce semantic overlap among these subspaces, while a remaining representation preserves semantic information that is not sufficiently explained by the prede- fined concepts. During candidate ranking, concept-specific similarities are activated according to the attributes detected in each query and are combined with the original dense embedding similarity through optimized non-negative weights.
Experiments on the Amazon Shopping Queries Dataset show that the proposed method achieves small but consistent improvements over the dense embedding baseline at different evaluation cutoffs. Further analysis indicates that the benefit of concept decomposition becomes more evident as the semantic structure of a query becomes more complex. Queries containing multiple simultaneous constraints obtain greater improvements than queries dominated by a single concept. These findings suggest that concept-specific subspaces provide complementary ranking evidence by reinforcing semantic constraints that may not be sufficiently distinguished within a single holistic embedding. The proposed framework therefore complements rather than replaces conventional embedding-based ranking and provides a more interpretable, con- trollable, and extensible approach to concept-aware query–product matching.
Chapter 1 Introduction 1
Chapter 2 Related Work 5
2.1 Dense Retrieval and Semantic Product Search 5
2.2 Multi-Vector and Multi-Field Retrieval Representations 6
2.3 Multi-Aspect, Attribute-Aware, and Concept-Level Decomposition 7
Chapter 3 Method 10
3.1 Overview 10
3.2 Shared Text Representation 11
3.3 Query-derived Concept Lexicon Construction 12
3.4 Concept Subspace Construction 13
3.5 Concept-guided Projection 13
3.6 Hierarchical Product Subspace 14
3.7 Remaining Semantic Representation 16
3.8 Concept-specific Similarity Computation 16
3.9 LLM-based Query Concept Detection 17
3.10 Attribute-Conditioned Score Aggregation 17
3.11 Final Ranking 18
Chapter 4 Experiments 20
4.1 Research Questions 20
4.2 Dataset 20
4.3 Task Definition 23
4.4 Evaluation Metrics 24
4.5 Baselines 25
4.5.1 Random Ranking 25
4.5.2 BM25 25
4.5.3 Official Benchmark Baseline 26
4.5.4 Dense Embedding Baseline 26
4.6 Experimental Settings 26
4.6.1 Implementation and Embedding Configuration 26
4.6.2 Concept Subspace Configuration 27
4.6.3 LLM-based Query Concept Detection and Subspace Activation 28
4.6.4 Weight Optimization Settings 29
4.6.5 Evaluation Procedure 30
4.7 Experiment 1: Overall Ranking Performance 31
4.8 Experiment 2: Effect of Query Concept Richness 34
4.9 Experiment 3: Improvement by Activated Subspace 35
4.10 Summary of Experimental Results 37
4.11 Qualitative Case Study 38
4.11.1 Case 1: Query 101087 39
4.11.2 Case 2: Query 46390 40
4.11.3 Case 3: Query 98030 41
Chapter 5 Discussion 43
5.1 Findings 43
5.1.1 Concept Decomposition Complements Dense Embeddings 43
5.1.2 Concept-Rich Queries Benefit More 43
5.1.3 Attribute-Specific Subspaces Provide Different Complementary Signals 44
5.1.4 Query Guidance Can Improve Search Interaction 45
5.2 Design Principles 46
5.2.1 Use Concept Signals as a Supplement to Holistic Similarity 46
5.2.2 Define Concepts According to the Decision Context 46
5.2.3 Prioritize Explicit and Comparable Concepts 47
5.2.4 Incorporate Hierarchical Concepts When Attribute Relationships Are Structured 47
5.2.5 Treat Weights as System-Level Design Parameters 47
5.3 Limitations and Future Work 48
Chapter 6 Conclusion 50
References 52
Bhattacharya, G., Kilari, N., Bhatia, A., and P., B. (2022). Datrnet: Disentangling fashion attribute embedding for substitute item retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1499–1508.
Freymuth, N., Liu, D., Ricatte, T., and Mansour, S. (2025). Hierarchical multi-field representations for two-stage e-commerce retrieval. arXiv preprint arXiv:2501.18707.
Gao, L., Dai, Z., and Callan, J. (2021). Coil: Revisit exact lexical match in information retrieval with contextualized inverted list. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3030–3042.
Google (2021). Google product taxonomy. [https://www.google.com/basepages/producttype/taxonomy.en-US.txt](https://www.google.com/basepages/producttype/taxonomy.en-US.txt). Version 2021-09-21.
Järvelin, K. and Kekäläinen, J. (2002). Cumulated gain-based evaluation of ir techniques. ACM Transactions on Information Systems, 20(4):422–446.
Jolliffe, I. T. and Cadima, J. (2016). Principal component analysis: A review and recent developments. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 374(2065):20150202.
Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781.
Khattab, O. and Zaharia, M. (2020). Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 39–48.
Kong, W., Khadanga, S., Li, C., Gupta, S. K., Zhang, M., Xu, W., and Bendersky, M. (2022). Multi-aspect dense retrieval. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 3178–3186.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. arXiv preprint arXiv:2005.11401.
Luo, C., Tang, X., Lu, H., Xie, Y., Liu, H., Dai, Z., Cui, L., Joshi, A., Nag, S., Li, Y., Li, Z., Goutam, R., Tang, J., Zhang, H., and He, Q. (2024). Exploring query understanding for amazon product search. arXiv preprint arXiv:2408.02215.
Nigam, P., Song, Y., Mohan, C., Meduri, V., Krishna, A., Srinivasan, A., Nair, K., Krishnamurthy, A., Elmagarmid, A., and Vasudev, B. (2019). Semantic product search. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 2876–2885.
OpenAI (2025). GPT-5 Mini Model. OpenAI API Documentation. Accessed: July 6, 2026.
Reddy, C. K., Màrquez, L., Valero, F., Rao, N., Zaragoza, H., Bandyopadhyay, S., Biswas, A., Xing, A., and Subbian, K. (2022). Shopping queries dataset: A large-scale esci benchmark for improving product search. arXiv preprint arXiv:2206.06588.
Robertson, S. E., Walker, S., Jones, S., Hancock-Beaulieu, M. M., and Gatford, M. (1994). Okapi at TREC-3. In Proceedings of the Third Text REtrieval Conference (TREC-3), pages 109–126.
Robertson, S. E. and Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4):333–389.
Salton, G. and Buckley, C. (1988). Term-weighting approaches in automatic text retrieval. Information Processing & Management, 24(5):513–523.
Storn, R. and Price, K. (1997). Differential evolution—a simple and efficient heuristic for global optimization over continuous spaces. Journal of Global Optimization, 11(4):341–359.
Wang, L., Yang, N., Huang, X., Yang, L., Majumder, R., and Wei, F. (2024). Multilingual e5 text embeddings: A technical report. arXiv preprint arXiv:2402.05672.
Xiong, L., Xiong, C., Li, Y., Tang, K., Liu, J., Bennett, P., Ahmed, J., and Overwijk, A. (2021). Approximate nearest neighbor negative contrastive learning for dense text retrieval. In 9th International Conference on Learning Representations (ICLR).
全文公開日期 2031/08/24