Graphazon is an applied research-focused project that combines query understanding, knowledge graph construction, and learning-to-rank models for Amazon-style product search. The system is powered byLLMs · Knowledge Graph Extraction · Agentic AI · Recommender Systems.
Product search is a complex problem that requires:
- Understanding user intent from ambiguous queries using LLM reasoning (LangChain)
- Mapping queries to relevant products, including substitutes and complements
- Leveraging structured and unstructured product information via a Knowledge Graph
- Improving ranking with graph-derived and semantic features
Graphazon addresses this by:
- Using LangChain for query understanding, intent detection, and attribute extraction.
- Building a task-specific Knowledge Graph from metadata and product text, including structured and LLM-assisted triples.
- Incorporating graph-derived features into a learning-to-rank model trained on real-world data.
- LangChain: LLM chains for query rewriting, intent detection, and attribute extraction.
- Knowledge Graphs (KGs): Structured + LLM-assisted KG construction to capture product relationships (Exact, Substitute, Complement).
- Embeddings & Features: Combine KG, product metadata, and query embeddings for ranking.
- Learning-to-Rank: Neural / gradient-boosted ranking models trained on ESCI labels.
- Pipelines: Modular orchestration for end-to-end experiments.
Graphazon uses the Amazon Shopping Queries Dataset (ESCI Benchmark):
- Provides queries, candidate products, and ESCI labels (Exact / Substitute / Complement / Irrelevant)
- Includes multilingual queries (English, Spanish, Japanese)
- Available as a git submodule in
data/shopping-queries-dataset/raw/amazon-science/esci-data/
data/
└── shopping-queries-dataset/
├── raw/
│ └── amazon-science/esci-data/ # Git submodule from amazon-science
└── processed/ # cleaned queries, products, KG triples, splits
To initialize the dataset:
git submodule update --init --recursiveGraphazon/
├── data/ # datasets
├── src/ # source code
│ ├── query_understanding/
│ ├── knowledge_graph/
│ ├── ranking/
│ ├── evaluation/
│ └── pipelines/
├── experiments/
├── notebooks/ # EDA, visualization, shows usage
├── scripts/ # utility scripts
├── pixi.toml # environment definition
└── README.md
-
query_understanding/: LLM chains and prompts for query parsing, intent detection, attribute extraction
-
knowledge_graph/: KG schema, structured and LLM-assisted extractors, graph feature generation
-
ranking/: Learning-to-rank models using features from KG and embeddings
-
evaluation/: Metrics, ablations, error analysis
-
pipelines/: Orchestration scripts that run end-to-end processes
Graphazon uses pixi for environment management.
pixi installThen initialize the ESCI dataset:
git submodule update --init --recursive
# Run query understanding
pixi run python -m src.pipelines.run_query_understanding
# Build the knowledge graph
pixi run python -m src.pipelines.build_kg
# Combine the features
pixi run python -m src.pipelines.build_features
# Train the learning-to-rank model
pixi run python -m src.pipelines.train_ranker
# Evaluate results
pixi run python -m src.pipelines.evaluate
Find examples with command-line arguments in scripts/
This project is for research and educational purposes. ESCI dataset is provided under its own license.