<!-- Source: https://docs.squirro.com/en/latest/technical/search/relevancy/scoring-profiles.html -->
# Scoring Profiles and Roles

Profiles: Project Creator, Search Engineer

This page presents an overview of _Scoring Profiles_ and _Scoring Roles_ within Squirro Cognitive Search.

Scoring profiles and roles are configured by search engineers, then used by project creators to finetune the search experience for end users.

> **Note**
>
> Optimizing each stage of the retrieval pipeline is crucial for achieving highly accurate search results.

[![Retrieval Stages](https://s3.amazonaws.com/download.squirro.net/docs/technical/search/relevancy/search-pipeline-stages.drawio.png)](https://s3.amazonaws.com/download.squirro.net/docs/technical/search/relevancy/search-pipeline-stages.drawio.png)

## Background: Document Relevancy

Squirro Cognitive Search uses a default scoring algorithm (BM25) to retrieve an initial _relevancy score_ indicating how relevant the document is for the given full-text search query. The _relevancy score_ is then used to rank documents from highly to partially relevant and ideally contains the right information in the top results.

However, relevancy is not a static concept. Within a specific project, relevancy may depend on the overall business objectives (use cases), user preferences, or other metrics.

For example, given two projects, in the first you might want to prioritize popular documents, whereas in the second you want to promote documents that have been recently modified.

In Squirro Cognitive Search, _Scoring Profiles_ and _Scoring Roles_ can be used to finetune relevancy in the following ways:

- **Scoring Profiles** define how additional ranking query clauses are built.
- **Scoring Roles** define what profiles should get applied based on the current user.

Reference: Learn more about [Document Relevancy](index.md#search-relevancy).

## Scoring Profiles

Project Configuration `topic.search.document-scoring-profiles`

Scoring profiles use document metadata as additional filtering criteria to return the most relevant documents according to the selected scoring profile.

Scoring profiles can either reference a configured profile from the project configuration by name or leverage a plugin without any project configuration required.

Reference: Learn more about [How to Use Scoring Profiles to Customize Document Relevancy Scoring](../how-to-guides/how-relevancy.md#search-how-relevancy).

**Out-of-the-Box Scoring Profiles**

Scoring profile plugins shipped out of the box include the following:

- QueryProfile
- ScriptProfile
- PluginProfile

**Using Scoring Profiles in Queries**

Scoring profiles can be referenced directly within search queries using the `profile:{}` syntax. There are two ways to reference profiles:

**Built-in Plugin Profiles** (no name parameter needed):

```
profile:{last_read}
profile:{semantic}
profile:{popular_item}
```

Learn more about [Scoring Profiles and Queries](../features/query-syntax.md#query-scoring-profiles) and [Examples Queries](scoring-plugins/index.md#scoring-plugins-example-queries).

**Custom Named Profiles** (name parameter required):

```
profile:{name:<PROFILE_NAME>}
```

> **Note**
>
> The `name:` prefix is required to explicitly use custom scoring profiles that are defined in the project settings. Without the `name:` prefix, Squirro only looks for installed scoring plugins by their name ([learn more](scoring-plugins/index.md#scoring-plugins-reference)).

**Using Custom Profiles in Search Queries**

Once defined, you can use these custom profiles in search queries. For example, searching for “machine learning” with custom boosting:

```
machine learning profile:{name:<PROFILE_NAME>}
```

### QueryProfile

This plugin formulates additional queries based on [Query Syntax](../features/query-syntax.md#search-query-syntax).

It’s useful for promoting documents that meet certain criteria.

For boolean conditions, all matching documents are equally boosted (see the _ScriptProfile_ for more sophisticated scoring approaches).

> **Note**
>
> The QueryProfile uses a Syntax Parser that combines statements using the `OR` operator (in contrast to the `AND` operator used to parse user queries)

This profile supports all feature that the Squirro Query Syntax offers, for example boosting fine-grained signals extracted from paragraphs or sentences during ingestion.

> **Note**
>
> It is also possible to configure Query Scoring Profiles that work with the user’s query terms, by using the {{query_terms}} syntax (Jinja Templating).

Reference: To learn more about using scoring profiles within Squirro query syntax, see [Scoring Profiles and Queries](../features/query-syntax.md#query-scoring-profiles).

#### Example Usage

The following example promotes items equally that have been modified within the last three months:

```
{
    "recently_modified__boost_equal": {
        "query": "$modified_at > now/d-3M/M"
    },
}
```

The following promotes items loaded from a specific source (_faq_) `OR` have been classified during data ingestion time (as being a _tutorial_, for example).
.. code-block:

```
{
    "from_knowledge_base": {
        "query": "source:faq is_tutorial:True^100"
    }
}
```

#### Advanced Usage: Relevance Ranking with Conditional Boosting

This example demonstrates how to boost documents more reliably when they meet
specific conditions:

- The document has `item_type:FAQ`.
- The **user’s query terms match** inside the `tags` field.

A stable combination of relevance signals is achieved using **multiplicative boosting**
via the `scale_by` function.

> **Note**
>
> The `scale_by` function multiplies the document’s initial relevancy score
> by the specified boost factor.
> In this example, documents that match both conditions will have their score computed as:
>
>
>
> `doc_score = initial_doc_score * 100 * 11`
>
>
>
> For more information on `scale_by` and other boosting techniques, see
> [Boosting Queries By Optional Ranking Signals](../features/query-syntax.md#query-level-boosting).

> **Note**
>
> The `{{query_terms}}` syntax is a placeholder and is dynamically replaced
> with the actual user query terms at runtime.

```json
{
    "boost_FAQ_and_matches_in_TAGS": {
        "query": "scale_by:{item_type:FAQ}^100 scale_by:{profile:{fulltext_match text:{{query_terms}} fields:tags}}^11"
    }
}
```

#### Advanced Usage: Personalization

Personalized query clause generation leverages user information through Jinja Templating.

Note: The templated information must be available via the User Service.

The following example boosts documents where the current user is one of the authors:

```
{
    "author_is_contributor": {
        "query": "author:{{user}}^100"
    }
}
```

The following example boosts documents that align with the user’s interests:

```
{
    "user_interest_aligns": {
        "query": "{% for interest in  interests%} tag:{{interest}}^10 {% endfor %}"
    }
}
```

### ScriptProfile

Using this plugin, you can formulate an [Elasticsearch Script Score Query](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-script-score-query.html) to implement your own custom scoring algorithm on top of the default search score.

This is useful for incorporating “static” signals that are independent of the query but highly correlated to relevance.

Example: You can promote previously modified items by applying a Gaussian Decay Function on the `modified_at` field.

The ScriptProfile plugin allows the highest flexibility to modify relevancy scoring, but comes with performance implications and should be used with caution.

The generated clause gets applied directly on the top-level [Squirro-Item](../../api/item-format.md#data-loading-item) and can only access common Fields and dynamically created Labels. There is no support for paragraph-level Signals/Entities.

#### Example Usage

The following example boosts documents that are more important in your domain based on pre-computed centrality scores like _PageRank_:

```json
"important_documents": {
    "script": {
        "source": "saturation(doc['kw_float']['pagerank'].value, 10)",
    },
    "debug": true
}
```

The following example boosts documents that have been recently modified considering recency (recent changes are more important - older ones less):

```json
"recently_modified__boost_decay": {
    "script": {
        "source": "decayDateGauss(params.origin, params.scale, params.offset, params.decay, doc['modified_at'].value)",
        "params": {
            "origin": "now",
            "scale": "30d",
            "offset" : "0",
            "decay" : 0.3
        }
    },
    "boost": 10,
    "debug": true
}
```

Note: The `origin` parameter supports date strings like `2022-08-01`, `2022-08-01T12:00:00Z` or `now` (current day).

### PluginProfile

This profile uses a custom ScoringPlugin that can leverage any kind of metadata from third-party systems to achieve higher document relevancy.
The built-in [PopularItem](#search-scoring-popular-items) plugin can be used with _PluginProfile_.

Custom Python extensions that implement the `RankClauseBuilder` interface (extensibility feature is currently under development).

They can leverage any kind of metadata from 3rd party systems to achieve higher document relevancy.

Plugin Profiles introduce the same flexibility to the generation of a Search Query (DSL) as [Pipelets](../../pipelets/index.md#pipelets) do for the Data Ingestion Pipeline.
Semantic search, for example, is implemented using a plugin that uses vector embeddings to find similar documents.

#### Example: Predefine how semantic search should be performed

Semantic search is usually executed by explicitly using the semantic ScoringPlugin with `profile:{ semantic }` ([learn more](scoring-plugins/retrieve.md#scoring-plugins-retrieve)) and custom arguments can be set to override default settings, like `profile:{semantic worker:snowflake}` ([learn more](scoring-plugins/index.md#scoring-plugins-reference)).
Scoring Profiles can also act as a shorthand to predefine custom settings, so that project admins/developers can simply use `profile:{ name:my_semantic }` and manage how semantic search should be performed from a central place.

```
{
    "my_semantic": {
        "asset_name": "semantic",
        "config": {
            "perform_only_knn": true,
            "vector_field": "embeddings.custom_snowflake_embedding_field",
            "worker": "snowflake-v2"
        }
}
```

#### Rank on Popular Items Example

This built-in plugin keeps track of the popularity of Items and adds additional boosting queries ad-hoc without relying on pre-computed popularity scores attached to the items.

The following example applies an additional boost on popular items:

```
{
    "popular_items": {
        "asset_name": "popular_item",
        "config": {
            "last_months": 3,

            # boost applies only if item was read at least 5 times
            "min_popularity": 5,

            # scope can be `project` or `user`
            "scope": "project"
        }
    }
}
```

### Plugin Reference:

To learn more, see the [Scoring Plugins](scoring-plugins/index.md#search-scoring-plugins) page.

## Scoring Roles

Project Configuration `topic.search.document-scoring-roles`

_Scoring Roles_ define what _Scoring Profiles_ should actually get executed.

Certain Profiles might make sense to get applied to all users, whereas others may only need to apply to a certain group of people.

Role configuration allows a versatile way of configuring the mapping between project users and scoring profiles.

Reference: Learn more about [How to Use Scoring Profiles to Customize Document Relevancy Scoring](../how-to-guides/how-relevancy.md#search-how-relevancy).

## Anatomy of a Search

A search query gets piped through the configured [Query Processing Workflow](../query-processing.md#query-processing) to preprocess the query before forwarding it to the Search Engine.

Common preprocessing tasks involve:

- Removal of unwanted terms like stopwords.
- Query classification: Language, Query Type (keyword vs. natural language question).
- Use-case-specific processing.

A processed example query might look like the following:

```
Original User Input
-------------------

    query:          "what are the annual reports of APPLE? $item_created_at > 2020"

Processed
---------

    user_terms:     "annual reports APPLE"
    user_filters:   "$item_created_at > 2020"
    query_type:     "question"
    query_language: "en"
```

The processed query information can then be combined with the configured Scoring Profiles to retrieve the most relevant document for the user.

_Figure 2: Combination of Query Term and Relevancy Signal Matching_

[![](https://s3.amazonaws.com/download.squirro.net/docs/technical/search/relevancy/search-pipeline_scoring_profile_overview_query_example.drawio.png)](https://s3.amazonaws.com/download.squirro.net/docs/technical/search/relevancy/search-pipeline_scoring_profile_overview_query_example.drawio.png)

## Profile Execution Stage: Re-Scoring

Not all profiles are suitable to be applied on all documents during the initial query phase, due to each additional ranking clause impacting latency.

Therefore the concept of **Profile Execution Stages** allows you to apply ranking profiles on either of the following:

- All relevant documents that match the overall user query (`stage: query`)
- Only on the most relevant subset of top N ranked documents that match the search query (`stage: rescore`).

_Figure 3: Applying Scoring Profiles on different stages._

[![](https://s3.amazonaws.com/download.squirro.net/docs/technical/search/relevancy/search-pipeline_scoring_profile_staging_query_example.drawio.png)](https://s3.amazonaws.com/download.squirro.net/docs/technical/search/relevancy/search-pipeline_scoring_profile_staging_query_example.drawio.png)

### Re-scoring: Precision meets Performance

#### Performance

Re-scoring is especially useful to improve precision by reordering just the top documents returned by the `query` phase, using a secondary (more expensive) algorithm, instead of applying the expensive algorithm to all matching documents.

This is important to consider when using the [Script Profile](#search-scoring-script-profile).

#### Precision

Furthermore, Re-scoring helps to combine global relevance information with query-centric relevance signals in a more meaningful way.

An example of this is adding PageRank scores to the final ranking only. [PageRank](https://en.wikipedia.org/wiki/PageRank) is a measure of the importance or informativeness of a document within a hyperlinked corpus of documents. Since the PageRank score is independent of a user query, it is a global feature of the corpus. The document relevance score (BM25), on the other hand, is dependent on a user query. Blindly combining the two scores (e.g., by multiplication) can easily result in one score overshadowing the other.

A more robust strategy is to use BM25 for coarse-grained selection of relevant documents in relation to the user query (recall), with subsequent re-evaluation of the top-scoring documents in relation to their PageRank score (improving precision).

_Apply rescoring on the pre-computed PageRank score on top ranked items only_

```json
"important_documents": {
    "script": {
        "source": "saturation(doc['kw_float']['pagerank'].value, 10)",
    },
    "debug": true,
    "stage": "rescore",
    "config": {
        "rescore": {
            "query_weight": 0.5,
            "rescore_query_weight": 5,
            "score_mode": "total",
            "window_size": 50
        }
    }
}
```

**Rescore Configuration Reference**

**`window_size`**

Type: int

Required: False

Default: 50

Control the number of top ranked documents that should be examined per shard.

**`query_weight`**

Type: float

Required: False

Default: 1.0

Control the relative importance of the original query.

**`rescore_query_weight`**

Type: float

Required: False

Default: 1.0

Control the relative importance of the applied rescore profile.

**`score_mode`**

Type: string

Required: False

Default: "total"

Control the way how the scores (original, rescore) are combined.

## Changelog

- [Squirro 3.6.1](../../../getting/release/3-6/squirro-3.6.x-release-notes/3.6.1-release-notes.md#release-notes-3-6-1): Initial Release of Scoring Profiles.
- [Squirro 3.6.2](../../../getting/release/3-6/squirro-3.6.x-release-notes/3.6.2-release-notes.md#release-notes-3-6-2): Added support for native Elasticsearch Scripts using Script Profile.
- [Squirro 3.6.3](../../../getting/release/3-6/squirro-3.6.x-release-notes/3.6.3-release-notes.md#release-notes-3-6-3): Introduced concept of Profile Execution-Stages (rescore vs. query).
