Skip to content

fix: support quantization ranges for int8 single-text embeddings - #11854

Closed
DebadityaHait wants to merge 1 commit into
deepset-ai:mainfrom
DebadityaHait:fix-int8-single-text-zero-embeddings
Closed

DebadityaHait wants to merge 1 commit into
deepset-ai:mainfrom
DebadityaHait:fix-int8-single-text-zero-embeddings

Conversation

@DebadityaHait

Copy link
Copy Markdown

Related Issues

Proposed Changes:

SentenceTransformersTextEmbedder with precision="int8"/"uint8" returns meaningless embeddings (all zeros or all equal values) and a downstream cosine division-by-zero in retrievers. Root cause: SentenceTransformer.encode calls quantize_embeddings(embeddings, precision) without ranges, so the min/max calibration is computed from the batch itself — degenerate for a single query text (min == max → step 0 → NaN → cast to int8).

  • SentenceTransformersTextEmbedder gains a quantization_ranges init parameter (shape (2, embedding_dim): min values in the first row, max in the second), serialized in to_dict/from_dict.
  • _SentenceTransformersEmbeddingBackend.embed handles a quantization_ranges kwarg: for int8/uint8 it encodes at float32 and quantizes via sentence_transformers.util.quantize_embeddings(..., ranges=...). Behavior without ranges is unchanged.
  • A warning is logged when a quantized precision is used without calibration ranges.
  • Release note added.

How did you test it?

  • Reproduced the bug: single text encoded with precision="int8" yields a degenerate embedding without ranges, correct distinct values with ranges.
  • New unit tests: backend quantization path (with/without ranges), component kwarg forwarding, warning, serialization round-trip.
  • New integration test test_run_quantization_with_ranges with sentence-transformers-testing/stsb-bert-tiny-safetensors asserting the embedding contains distinct int values.
  • hatch run test:unit test/components/embedders/ → 167 passed; hatch run test:types → clean; hatch run fmt → clean; pre-commit hooks pass.

Notes for the reviewer

  • Only ranges is exposed; quantize_embeddings also accepts calibration_embeddings, which could be added later if users prefer passing raw embeddings from their Document Store.
  • Scope kept to the text embedder (single-text batches are where the calibration degenerates); SentenceTransformersDocumentEmbedder behavior is unchanged, though the backend support would allow extending it there too.
  • This PR was generated with AI assistance (per the contributing guide's AI disclosure requirement).

Checklist

  • I have read the contributors guidelines and the code of conduct.
  • I have updated the related issue with new insights and changes.
  • I have added unit tests and updated the docstrings.
  • I've used one of the conventional commit types for my PR title: fix:, feat:, build:, chore:, ci:, docs:, style:, refactor:, perf:, test: and added ! in case the PR includes breaking changes.
  • I have documented my code.
  • I have added a release note file, following the contributors guidelines.
  • I have run pre-commit hooks and fixed any issue.

@DebadityaHait
DebadityaHait requested a review from a team as a code owner July 2, 2026 20:16
@DebadityaHait
DebadityaHait requested review from bogdankostic and removed request for a team July 2, 2026 20:16
@vercel

vercel Bot commented Jul 2, 2026

Copy link
Copy Markdown
Contributor

@DebadityaHait is attempting to deploy a commit to the deepset Team on Vercel.

A member of the Team first needs to authorize it.

@CLAassistant

CLAassistant commented Jul 2, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@sjrl

sjrl commented Jul 3, 2026

Copy link
Copy Markdown
Contributor

Hey @DebadityaHait thanks for opening the PR! The sentence transformer components have been moved to our core integrations repo here https://github.com/deepset-ai/haystack-core-integrations/tree/main/integrations/sentence_transformers so please close this PR and open the fix our core integrations repo.

The sentence transformer components in this repo are deprecated and will soon be removed in Haystack v3.

@DebadityaHait

Copy link
Copy Markdown
Author

Thanks for clarifying. I’ll close this PR and port the fix to haystack-core-integrations.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SentenceTransformersTextEmbedder zero embeddings for single text queries when using precision="int8"

3 participants