PROJECT METHODOLOGY

What the explorer shows.

EntityLens illustrates business entity resolution with records and their provided ground-truth labels. It does not execute a matching model or report model accuracy.

Reference businesses and matching records

Source 1 is the reference source. Each reference business may match zero, one, or many records from Sources 2 and 3. A target is a positive example when its ID appears in the reference business’s ground-truth match list. A negative example is a target absent from that complete list. A business with an empty list is a singleton.

Example data and selection

The bundled snapshot contains synthetic records: 100 reference businesses, 160 true links, and 20 singletons. All true links for each displayed business are preserved. Up to four negative examples from each target source are selected for similar names, with same-country examples prioritized. These are comparison examples, not a pipeline candidate set.

The exporter can replace this snapshot with provided challenge training data. It selects reference records deterministically and considers a bounded target sample alongside all required positive records. The website always identifies whether its snapshot is synthetic or training data.

Text agreement and ground truth

The pair inspector compares names, addresses, and country. Different words are highlighted. “Same normalized text” means the two fields agree after Unicode normalization, case folding, and whitespace handling. Missing fields are shown explicitly and do not establish a match. These descriptions are text comparisons, not confidence scores.

The project’s matching pipeline

Read the complete architecture for the inference flow diagram, model defaults, training splits, calibration policy, output contract, storage, and deployment.

The runnable Python pipeline preserves raw Unicode, builds lexical views, and generates candidates independently using a trained Indic BGE-M3 bi-encoder and character-trigram TF-IDF. It unions those candidates and derives field-level pair features. XGBoost handles confident decisions; uncertain pairs are routed to a supervised BGE multilingual cross-encoder.

Training separates fitting, development, calibration, and audit groups. Calibrated thresholds optimize macro F0.5, including singletons. Predictions allow multiple matches and must remain within the generated candidate set. The architecture walkthrough explains this methodology; it is not a trace of a model run.

Scale and evaluation

The challenge’s approximately 24.2 million source records and 1.73 million Source 1 test queries describe the challenge dataset. They are separate from the displayed demo snapshot. Trained-model scores require a measured evaluation with its split, sample size, and target-pool scope.

Source code and reproducibility

View the project repository for the pipeline, training configuration, preserved documentation, and website export instructions.