Multilingual records prepared for semantic search while preserving source identity.
Each of the 43 records contains vector_text_en and vector_text_fa: normalized source strings that may be supplied to an embedding model. Snapshot v1.0.0 publishes no numeric vector, dimension, model identifier, similarity score or semantic index.
Every source string remains attached to its record ID; no claim is made about future embedding identity.
Persian and English source strings are stored separately in the same record.
Any later vector search must name its model, distance measure, corpus version and evaluation set.