Sentence Transformers
So i was reading about the latent space of sentence transformer. My initial hypothesis was that data was distrubuted uniformly in the latent space. But what i discovered was data was organized in cone structure.
Algorithm
- so what happens is that when we pass in a sentence of say 6 words.
- first it create embeddings of each word at layer 0.
- these are just raw embeddigns, no causal relation ships
- Then that 6 vectors are passed through layer 1.
- Now layer 1 would update the relations ships based on its learned mechanisms
- That 6 vectors would keep updating untill the end of transformer with 13 layers. at the 13th layer we would get the complete nuanced vectors.
© 2026 bsybin