For best job speed and to avoid memory issues, use aggregated signals instead of raw signals as input for this job.
Input
This job takes one or more of the following as input:Signal data
This input is required; additional input is optional. Signal data can be either raw or aggregated. The job runs faster using aggregated signals. When raw signals are used as input, this job performs the aggregation. Use thetrainingCollection/Input Collection parameter to specify the collection that contains the signal data.
Misspelling job results
Token and Phrase Spell Correction job results can be used to avoid finding mainly misspellings, or mixing synonyms with misspellings. Use themisspellingCollection/Misspelling Job Result Collection parameter to specify the collection that contains these results.
Phrase detection job results
Phrase Extraction job results can be used to find synonyms with multiple tokens, such as “lithium ion” and “ion battery”. Use thekeyPhraseCollection/Phrase Extraction Job Result Collection parameter to specify the collection that contains these results.
Keywords
A keywords list in the blob store can serve as a blacklist to prevent common attributes from being identified as potential synonyms. The list can include common attributes such as color, brand, material, and so on. For example, by including color attributes you can prevent “red” and “blue” from being identified as synonyms due to their appearance in similar queries such as “red bike” and “blue bike”. The keywords file is in CSV format with two fields:keyword and type. You can add your custom keywords list here with the type value “stopwords”. An example file is shown below:
keywordsBlobName/Keywords Blob Store parameter to specify the name of the blob that contains this list.
Custom Synonyms
For some deployments there might be a need to use existing synonym definitions. You can import existing synonyms into the Synonym and Similar Queries Detection job as a text file. Upload your synonyms text file to the blob store and reference that file when creating the job.Output
The output collection contains two tables distinguished by thedoc_type field.
The similar queries table
Ifquery leads to clicks on documents 1, 2, 3, and 4, and similar_query leads to clicks on documents 2, 3, 4, and 5, then there is sufficient overlap between the two queries to consider them similar.
A statistic is constructed to compute similarities based on overlap counts and query counts. The resulting table consists of documents whose doc_type value is “query_rewrite” and type value is “simq”.
The similar queries table contains similar query pairs with these fields:
The synonyms table
The synonyms table consists of documents whosedoc_type value is “query_rewrite” and type value is “synonym”: