A fast, parallel Rust tool for namespace-based summarization of massive RDF graphs — extracts structure, infers namespaces, normalizes triples, and generates interactive visualizations.
- Background
- Install
- Quick Start
- Usage
- Outputs & Visualization
- Auxiliary Commands
- Validation
- Citation
- Contributing
- License
RDF graphs at web scale (hundreds of millions to billions of triples) are difficult to explore and understand. chilon_rs addresses this by:
- Namespace Inference — automatically discovers namespace prefixes from IRIs using a community prefix table and statistical segmentation.
- Triple Normalization — groups triples by inferred namespace, counts occurrences, and filters by minimum frequency to produce a compact summary.
- Visualization — emits JSON data and a self-contained HTML/JS visualization for interactive exploration of the summary.
The algorithm is based on: dos Santos & Leal (2023). "Summarization of Massive RDF Graphs Using Identifier Classification." ICCS 2023.
Requires Rust 1.77+.
git clone https://github.com/andrefs/chilon_rs.git
cd chilon_rs
cargo build --releaseThe binary will be at cli/target/release/chilon_rs.
cargo install chilon_rs# Build
cargo build --release
# Process an RDF file (Turtle format)
./cli/target/release/chilon_rs mygraph.ttlThis creates a dated folder under results/YYYYMMDD/ containing:
- Normalized triples grouped by namespace
- Namespace prefix table
summary.json+visualization.htmlfor interactive browsing
# Process a single file with defaults (namespace inference ON, strict mode)
chilon_rs graph.ttl| Flag | Description |
|---|---|
--no-infer-ns |
Disable namespace inference; use only the built-in community prefix table. |
-i, --ignore-unknown |
Skip triples whose namespace cannot be resolved (instead of erroring). |
-h, --help |
Show help message. |
-V, --version |
Print version. |
# Process multiple files in one run
chilon_rs file1.ttl file2.ttl file3.ttlWorker threads are auto-tuned: max(2, min(files + 1, CPU cores - 2)).
After processing, a folder results/YYYYMMDD-N/ is created with:
| File | Description |
|---|---|
output.ttl |
Normalized triples grouped by namespace (Turtle). |
vis-data.json |
Visualization data (nodes, edges, aliases). |
namespaces.tsv |
Inferred + community prefix table (prefix, URI, count). |
normalized.tsv |
Normalized triples: subject\tpredicate\tobject\tcount. |
tasks.json |
Processing metadata (timings, triple counts, stages). |
chilon.log |
Detailed execution log. |
The visualization is a JavaScript module that must be served over HTTP (browsers block ES modules on file://). Two steps:
-
Build the visualization assets (requires Node.js ≥ 18 + yarn):
cargo run --bin gen-viz -- results/20260909-9
This runs
vite buildinsidechilon-viz/and copiesdist/into the results folder. -
Serve the
dist/folder and open in browser:cd results/20260909-9/dist && python3 -m http.server 8000 # Then open http://localhost:8000
Two additional binaries ship with the crate:
# Generate visualization from an existing results folder
cargo run --bin gen-viz -- results/20260909-9
# Quick test/debug: parse RDF and show basic stats without full pipeline
cargo run --bin test-files -- file1.ttl file2.ttlchilon_rs has been validated on 11 real-world RDF graphs spanning from a few MB / <1M triples to 90+ GB / billions of triples:
| Dataset | Domain |
|---|---|
| ClaimsKG | Fact-checking claims |
| CrunchBase | Company/startup data |
| DbKwik | Wikipedia infobox extraction |
| DBLP | Bibliography / publications |
| DBpedia | Wikipedia structured data |
| KBpedia | Knowledge base ontology |
| LinkedMDB | Movie database |
| OpenCyc | Common-sense knowledge base |
| Wikidata | General knowledge graph |
| WordNet | Lexical database |
| YAGO | Ontology from Wikipedia |
Summaries and visualizations: https://andrefs.github.io/chilon_rs
If you use chilon_rs or the underlying algorithm in your work, please cite:
dos Santos, A.F., Leal, J.P. (2023). Summarization of Massive RDF Graphs Using Identifier Classification. In: Ojeda-Aciego, M., Sauerwald, K., Jäschke, R. (eds) Graph-Based Representation and Reasoning. ICCS 2023. Lecture Notes in Computer Science. Springer, Cham. https://doi.org/10.1007/978-3-031-40960-8_8
The 11 corpora used for validation are publicly available:
Contributions are welcome! Please see CONTRIBUTING.md for guidelines on:
- Setting up the development environment (
cargo test,cargo clippy --all-targets -- -D warnings) - Code style and commit conventions
- Opening issues and pull requests
Released under the MIT License — see LICENSE for details.
Copyright (c) 2023 Alexandre F. dos Santos