How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code

2026-08-26 · Hugging Face

Background: Search as the foundation of the Papers with Code revival

In the Papers with Code revival, the team’s goal is to make open AI research easier to access and understand. That means users should quickly find papers, related artifacts, and state-of-the-art results across domains. Search is therefore a core platform capability, used both on the website and through the `pwc` search CLI/Skill for agents.

The post stresses that paper search is different from generic text search. A useful system must:

  • match exact titles and arXiv IDs,
  • understand semantic queries even when wording differs,
  • handle navigational intent (e.g., “the original BERT paper”),
  • tolerate incomplete titles and typos,
  • and remain responsive when model services are cold or temporarily unhealthy.

Why hybrid retrieval

Papers with Code uses a hybrid search design rather than a single retrieval method:

  • Lexical retrieval via PostgreSQL full-text search for fast exact matching,
  • Semantic retrieval via pgvector for dense embedding recall,
  • Rank fusion via Reciprocal Rank Fusion (RRF) to combine both result sets.

Based on prior RAG experience, the team notes that hybrid retrieval typically outperforms keyword-only or vector-only systems because it captures both precision and semantic similarity. They also mention rerankers (cross-encoders) can improve quality further, but at additional latency and compute cost.

How Hugging Face services are split by responsibility

Three Hugging Face services support dense retrieval operations:

  • Hugging Face Jobs for burstable GPU batch embedding of the corpus,
  • Hugging Face Storage Buckets for durable handoff between DB exports, experiments, and Jobs,
  • Hugging Face Inference Endpoints for low-latency embeddings in live queries and incremental updates.

The system currently maintains embeddings for more than 110,000 papers from arXiv and Daily Papers.

Architectural decision: separate offline build from online serving

The team intentionally separates expensive offline processing from online request handling:

1. Offline corpus build runs throughput-oriented embedding jobs.

2. Online search path keeps only a small query-embedding step behind a protected endpoint.

If the endpoint is cold, overloaded, or unhealthy, search immediately falls back to full-text retrieval. This keeps the experience fast and resilient while preserving semantic capabilities when available.

A strict, versioned embedding contract

A major lesson is that embedding pipelines often fail subtly: model revisions drift, query/document prompts get mixed, vector dimensions change, or updated abstracts no longer match stored vectors.

To prevent this, they treat embedding format as a versioned API. Each paper is encoded as:

`normalized title + "\n\n" + normalized abstract`

For every vector generation, they record:

  • model repository and exact revision,
  • output dimension,
  • input-format version,
  • whether input is query or document,
  • normalization method,
  • content hash of source title and abstract.

This contract ensures traceability from export and GPU inference through PostgreSQL storage to online retrieval.

Production model choices

Production generation uses Qwen/Qwen3-Embedding-0.6B, pinned to an exact revision, with 256-dimensional L2-normalized vectors. Model selection was informed by the MTEB leaderboard.

The post highlights two newer embedding-model features used in design decisions:

  • Dynamic embedding size (MRL / Matryoshka Representation Learning) to trade quality against speed and storage,
  • Instruction prompts with separate document and query prompts.

They selected 256 dimensions to keep search fast.

Batch pipeline details

Full-corpus embedding is treated as a classic batch workload: short-lived GPU need, high throughput, and no idle resource consumption between runs. Hugging Face Jobs fits this pattern via command + hardware flavor (+ optional Docker image).

The corpus build process:

  • exports latest paper versions from a repeatable-read PostgreSQL snapshot,
  • streams rows instead of loading the full catalog in memory,
  • writes bounded JSONL shards,
  • produces a manifest with row counts and SHA-256 checksums,
  • syncs immutable run artifacts to a private Storage Bucket,
  • mounts the bucket with `hf-mount` into an `l4x1` Job (NVIDIA L4, 24GB VRAM).

Overall, the design emphasizes reproducibility, reliability, and low-latency serving while scaling semantic search in production.

Source