Skip to content

Add native reranking API for cross-encoder models #323

Description

@leehack

Motivation

llama.cpp has mature reranking support, including rank pooling, classifier metadata, named rerank chat templates, and server /rerank endpoints. The pinned b10333 native bindings already expose the required llama.cpp functions, but llamadart currently only exposes embedding-based similarity ranking.

Upstream: ggml-org/llama.cpp#9510

Proposed scope

  • Add an explicit backend reranking capability.
  • Add a public LlamaEngine.rerank(query, documents) API with typed ranked results.
  • Use the model's named rerank template when available, with a compatible fallback prompt construction.
  • Support batched document scoring with distinct sequence IDs.
  • Handle both single-score and multi-class classifier output.
  • Fail with a typed, actionable error for incompatible models/backends.
  • Document supported platforms, expected model types, and scoring semantics.

This should initially be implemented in llamadart; b10333 appears to expose the required native symbols without a new llamadart-native release.

Validation

  • Compare ordering and scores with the same model and inputs through upstream llama-server.
  • Cover single-document and batched reranking.
  • Cover multi-class classifier metadata where available.
  • Cover long input/context-limit failure.
  • Cover non-reranker and unsupported-backend rejection.
  • Run a real BGE reranker smoke on representative native platforms.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority:P2Planned next: useful unblocked work or validation after P1 items

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions