Using KV cache as embeddings

2 points | by mxudev an hour ago

1 comments

  • mxudev an hour ago

    What if instead of one giant vector as embedding, we use multiple (K, V) pairs. In this work we demonstrated this is feasible, and got reranker behavior at retrieval cost (without extra backbone pass)