Skip to content

User Guide

Row2Vec turns rows into vectors. You give it a DataFrame; it gives you back a numeric matrix with one row per input row, having dealt with the mixed dtypes, the missing values, and the scaling on the way.

Almost everything you want to do with a table downstream — cluster it, find the nearest neighbours of a record, plot it, feed it to a model that only speaks floats — assumes the rows are already points in a metric space. This guide covers the concepts, helps you pick a method, and works through the common tasks. Every code block on this page is executed by the test suite, so the examples stay correct against the installed version.

New here? Read Core concepts and Choosing a method, then jump to the section closest to your problem.


Core concepts

One call for every method

Every embedding is produced by learn_embedding. You choose the method with mode; everything else about the call stays the same.

import row2vec

df = row2vec.generate_synthetic_data(120)
embeddings = row2vec.learn_embedding(df, mode="pca", embedding_dim=3)

assert embeddings.shape == (120, 3)

Swapping methods means swapping one argument — the call site never changes.

The output is a DataFrame aligned with your input

The result has the same length and index as the input, so you can put it straight back next to the original columns.

import pandas as pd

import row2vec

df = row2vec.generate_synthetic_data(50)
embeddings = row2vec.learn_embedding(df, mode="pca", embedding_dim=2)

assert list(embeddings.index) == list(df.index)
combined = pd.concat([df, embeddings], axis=1)
assert len(combined.columns) == len(df.columns) + 2

Mixed dtypes are handled for you

The input can hold numbers, categories, and gaps at once. Numeric columns are scaled, categorical columns are encoded, and missing values are imputed — you do not have to prepare any of it.

import numpy as np
import pandas as pd

import row2vec

df = pd.DataFrame(
    {
        "amount": [10.0, 240.0, 35.5, np.nan, 88.0, 12.0, 300.0, 45.0],
        "region": ["north", "south", "north", "east", "south", "east", "north", "south"],
        "tier": ["a", "b", "a", "b", "a", "b", "a", "b"],
    }
)

embeddings = row2vec.learn_embedding(df, mode="pca", embedding_dim=2)

assert embeddings.shape == (8, 2)
assert not embeddings.isna().to_numpy().any()  # no NaN survives into the output

embedding_dim is the width of the output

It is the number of columns you get back — how much room the method has to describe a row. Two or three for plotting; five to fifty as features for another model.

import row2vec

df = row2vec.generate_synthetic_data(60)

for dim in (2, 5):
    assert row2vec.learn_embedding(df, mode="pca", embedding_dim=dim).shape[1] == dim

If you would rather not pick, see Choosing the dimension automatically.


Choosing a method

Start from what you want the vectors for:

Your situation Mode Key parameter
A fast, interpretable baseline pca embedding_dim
The structure is non-linear unsupervised hidden_units, max_epochs
A 2-D picture showing clusters tsne perplexity
A 2-D picture that also keeps global layout umap n_neighbors, min_dist
One vector per category, not per row target reference_column
You know which rows are alike contrastive auto_pairs, margin

By data characteristics:

Rows Structure Good default
Any Unknown — you are exploring pca first, then unsupervised
Thousands+ Non-linear, plenty of data to fit unsupervised
Up to a few thousand You want to see it umap (or tsne)
Any You have labelled pairs or a grouping column contrastive

When-to-use, in one line each:

  • pca — linear, deterministic, and instant. Always worth running first: if a few components already separate what you care about, stop here.
  • unsupervised — an autoencoder. Captures interactions PCA cannot, at the cost of training time and hyperparameters. Wants a few thousand rows.
  • tsne — for visualisation only. Excellent at revealing clusters, but distances between clusters are not meaningful, and it cannot embed new rows.
  • umap — usually the better plot: faster than t-SNE and keeps more of the global arrangement.
  • target — flips the question around: instead of embedding rows, embed the values of one column by the rows they occur in.
  • contrastive — supervised by pairs. Use it when you know that certain rows should be close (same customer, same cluster, same label).

Comparing modes on your own data

The tables above are rules of thumb. To see which mode suits your table, ask:

import row2vec

df = row2vec.generate_synthetic_data(150)

report = row2vec.compare_modes(df, target="Country", modes=["pca", "tsne"])

assert report.loc["pca", "status"] == "ok"
assert 0.0 < report.loc["pca", "trustworthiness"] <= 1.0
print(report[["trustworthiness", "downstream_score", "fit_seconds"]])

compare_modes fits each mode on the same training rows and scores it on rows it did not see, so a mode that memorises its training data gains nothing:

  • trustworthiness — whether a row's nearest neighbours in the embedding are its nearest neighbours in the preprocessed table. 1.0 means the neighbourhood structure survived.
  • downstream_score — given a target column, how well a k-nearest-neighbour model on the embedding predicts it: accuracy for a categorical target, R² for a numeric one. The target column is withheld from every mode as an input (only mode="target" uses it, as its label).
  • A baseline row gives the same score on the preprocessed features with no embedding, which is the number an embedding has to justify itself against.

t-SNE cannot embed unseen rows, so it is scored on trustworthiness alone, over a sample of at most tsne_max_rows rows (its cost grows steeply with row count). Modes that need TensorFlow show up as unavailable when it is not installed, and a mode that raises is reported as failed with the error, so one bad mode never hides the rest.

On the bundled data (examples/compare_modes_real_data.py, 4 dimensions, one run), PCA's 4 numbers predict Titanic survival about as well as the full preprocessed table (0.80 against 0.82 accuracy), and the supervised target mode did best (0.86). Treat figures like these as an illustration of the report, not a ranking: they move with the data, the split and the seed. Pass max_epochs=20 (or any other learn_embedding argument) to keep the neural modes quick.


The methods

PCA

The linear baseline. Fast, deterministic, and its components come out ordered by how much variance they account for.

import row2vec

df = row2vec.generate_synthetic_data(200)
embeddings = row2vec.learn_embedding(df, mode="pca", embedding_dim=3)

assert embeddings.shape == (200, 3)
# Ordered by explained variance, so the first component is the widest.
assert embeddings.iloc[:, 0].var() >= embeddings.iloc[:, 2].var()

Autoencoder (unsupervised)

A neural network trained to reconstruct each row through a narrow bottleneck; the bottleneck is the embedding. hidden_units sets the layers before it — a single integer for one layer, a list for several.

import row2vec

df = row2vec.generate_synthetic_data(200)
embeddings = row2vec.learn_embedding(
    df,
    mode="unsupervised",
    embedding_dim=4,
    hidden_units=[32, 16],  # two hidden layers
    max_epochs=3,  # kept small for this example; use far more in practice
    verbose=False,
)

assert embeddings.shape == (200, 4)

max_epochs bounds training; with early_stopping=True (the default) it stops sooner once the reconstruction loss plateaus.

t-SNE and UMAP

Both exist to be looked at. perplexity (t-SNE) and n_neighbors (UMAP) control how large a neighbourhood each point is fitted against — smaller values favour tight local structure, larger ones a smoother global picture.

import row2vec

df = row2vec.generate_synthetic_data(150)

tsne = row2vec.learn_embedding(df, mode="tsne", embedding_dim=2, perplexity=10)
umap = row2vec.learn_embedding(df, mode="umap", embedding_dim=2, n_neighbors=10, min_dist=0.1)

assert tsne.shape == umap.shape == (150, 2)

perplexity must be small relative to the number of rows; Row2Vec says so directly rather than letting scikit-learn fail obscurely.

import pytest

import row2vec

df = row2vec.generate_synthetic_data(30)

with pytest.raises(ValueError, match="should be less than"):
    row2vec.learn_embedding(df, mode="tsne", embedding_dim=2, perplexity=100)

Target-based embeddings

Supervise the encoder with a label column: rows sharing a value are pushed together in the embedding space. The result is one vector per row, like every other mode, and aggregate_by_reference=True collapses it to one vector per distinct value — useful for turning a high-cardinality categorical into a small dense feature.

import row2vec

df = row2vec.generate_synthetic_data(200)
row_vectors = row2vec.learn_embedding(
    df,
    mode="target",
    reference_column="Country",
    embedding_dim=2,
    max_epochs=3,
    verbose=False,
)

# Since 0.4.0 target mode returns one row per input row, carrying df's index,
# so it joins straight back on like every other mode.
assert len(row_vectors) == len(df)
assert list(row_vectors.index) == list(df.index)

# For one row per distinct country, ask for it explicitly.
country_vectors = row2vec.learn_embedding(
    df,
    mode="target",
    reference_column="Country",
    embedding_dim=2,
    max_epochs=3,
    verbose=False,
    aggregate_by_reference=True,
)
assert len(country_vectors) == df["Country"].nunique()

Contrastive embeddings

Supervision by example: tell the model which rows should end up close together and which should not.

import row2vec

df = row2vec.generate_synthetic_data(120)

embeddings = row2vec.learn_embedding(
    df,
    mode="contrastive",
    embedding_dim=3,
    similar_pairs=[(0, 1), (2, 3)],
    dissimilar_pairs=[(0, 50), (1, 60)],
    contrastive_loss="contrastive",
    max_epochs=3,
    verbose=False,
)

assert embeddings.shape == (120, 3)

If you don't have pairs to hand, auto_pairs derives them: "categorical" (rows sharing a category value are alike), "cluster", "neighbors", or "random".

import row2vec

df = row2vec.generate_synthetic_data(120)

embeddings = row2vec.learn_embedding(
    df,
    mode="contrastive",
    embedding_dim=2,
    auto_pairs="categorical",
    reference_column="Country",
    max_epochs=3,
    verbose=False,
)

assert embeddings.shape == (120, 2)

Preparing data

Missing values

AdaptiveImputer analyses the missingness of each column and picks a strategy to match. learn_embedding runs it for you, but you can use it on its own.

import numpy as np
import pandas as pd

from row2vec import AdaptiveImputer, ImputationConfig, MissingPatternAnalyzer

df = pd.DataFrame(
    {
        "age": [25.0, np.nan, 41.0, 33.0, np.nan, 29.0],
        "city": ["rome", "oslo", None, "rome", "oslo", "rome"],
    }
)

analysis = MissingPatternAnalyzer(ImputationConfig()).analyze(df)
assert analysis["total_missing"] == 3

imputed = AdaptiveImputer(ImputationConfig()).fit_transform(df)
assert imputed.isna().sum().sum() == 0
assert len(imputed) == len(df)

Choose the trade-off explicitly when you care: numeric_strategy="mean" with prefer_speed=True for quick iteration, numeric_strategy="knn" when accuracy matters more than time.

import numpy as np
import pandas as pd

from row2vec import AdaptiveImputer, ImputationConfig

df = pd.DataFrame({"x": [1.0, 2.0, np.nan, 4.0, 5.0, np.nan, 7.0, 8.0]})

fast = AdaptiveImputer(ImputationConfig(numeric_strategy="mean", prefer_speed=True)).fit_transform(
    df
)
accurate = AdaptiveImputer(ImputationConfig(numeric_strategy="knn", knn_neighbors=3)).fit_transform(
    df
)

assert fast.isna().sum().sum() == 0
assert accurate.isna().sum().sum() == 0

Set preserve_missing_patterns=True to keep the fact that a value was missing as its own feature — often predictive in itself.

Categorical columns

The encoder picks a strategy from the column's cardinality: one-hot for a handful of values, ordinal or target encoding as the count grows, learned entity embeddings for the largest. You can inspect that decision:

import pandas as pd

from row2vec import CategoricalAnalyzer, CategoricalEncodingConfig

df = pd.DataFrame({"colour": ["red", "green", "blue", "red", "green", "blue"]})

analysis = CategoricalAnalyzer(CategoricalEncodingConfig()).analyze_column(df["colour"])

assert analysis["cardinality"] == 3
assert analysis["recommended_strategy"]  # a strategy name, e.g. "onehot"

Datetime and boolean columns

A datetime column becomes cyclical features plus a trend. For each of hour, weekday, day-of-month and month that varies in the training data, the encoder emits a sine/cosine pair, so 23:00 sits next to 00:00 and December next to January. One standardised "elapsed time" feature separates 2019 from 2024. Missing timestamps take the training median, and timezone-aware columns are converted to UTC. Boolean columns (including nullable boolean) become 0/1.

import pandas as pd

from row2vec import learn_embedding

df = pd.DataFrame(
    {
        "placed_at": pd.date_range("2024-01-01", periods=40, freq="13h"),
        "express": [i % 3 == 0 for i in range(40)],
        "amount": [float(i % 7) for i in range(40)],
    }
)

embeddings = learn_embedding(df, mode="pca", embedding_dim=2, enable_logging=False)
assert embeddings.shape == (40, 2)

Text columns

A string column is a category unless you say otherwise. To treat it as free text, name it in text_columns. The default encoding is TF-IDF followed by a truncated SVD, which needs no extra dependency; text_dim is the most features each column gets (fewer if the vocabulary is smaller).

import pandas as pd

from row2vec import EmbeddingConfig, PreprocessingConfig, learn_embedding

df = pd.DataFrame(
    {
        "price": [float(i) for i in range(30)],
        "review": ["great apple pie", "blue sky today", "apple tart recipe"] * 10,
    }
)

# `config` carries the preprocessing settings; mode and size are keywords.
config = EmbeddingConfig(
    preprocessing=PreprocessingConfig(text_columns=["review"], text_dim=4),
)
embeddings = learn_embedding(df, mode="pca", embedding_dim=2, config=config, enable_logging=False)
assert embeddings.shape == (30, 2)

To use a sentence-embedding model instead, pass text_encoder: any callable that takes a list of strings and returns an array of shape (len(texts), dim). Each text column goes through it independently, in batches, and its output width is fixed when the model is fitted.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")
encode = lambda texts: model.encode(texts, normalize_embeddings=True)

config = EmbeddingConfig(
    preprocessing=PreprocessingConfig(text_columns=["review"], text_encoder=encode)
)
embeddings = learn_embedding(df, mode="pca", embedding_dim=8, config=config)

Text features are not rescaled, so a hook that returns unit-length vectors gives each text column roughly the weight of one standardised numeric column.

A custom text_encoder is code, and a saved model never runs code from its file, so the function is not saved. The model records that it needs one, and loading it again requires the same function:

model = row2vec.load_model("reviews.r2v", text_encoder=encode)

Without it load_model raises a ModelFormatError that says so. The default TF-IDF encoder is saved with the model and needs nothing at load time. Only text_columns and text_dim can be set in a YAML config; the callable cannot.

Scaling the output

scale_method rescales the embedding after it is computed — useful when a downstream model expects a bounded range.

import row2vec

df = row2vec.generate_synthetic_data(100)

scaled = row2vec.learn_embedding(
    df,
    mode="pca",
    embedding_dim=2,
    scale_method="minmax",
    scale_range=(0.0, 1.0),
)

values = scaled.to_numpy()
assert values.min() >= -1e-6
assert values.max() <= 1.0 + 1e-6

The options are "minmax", "standard", "l2", "tanh", and "none".


Tuning

Choosing the dimension automatically

auto_select_dimension evaluates candidate widths and recommends one, so you do not have to guess.

import row2vec

df = row2vec.generate_synthetic_data(150)
recommended_dim, details = row2vec.auto_select_dimension(
    df, methods=["pca_variance"], max_dimension=5
)

assert 1 <= recommended_dim <= 5
assert "method_results" in details

Searching the architecture

For the neural modes, search_architecture explores layer counts, widths, dropout rates, and activations, and returns the best configuration it found.

import row2vec
from row2vec import ArchitectureSearchConfig, EmbeddingConfig, NeuralConfig

df = row2vec.generate_synthetic_data(150)

base_config = EmbeddingConfig(
    mode="unsupervised", embedding_dim=3, neural=NeuralConfig(max_epochs=2)
)
search_config = ArchitectureSearchConfig(
    method="random", max_trials=2, verbose=False, layer_range=(1, 2)
)

best_architecture, _result = row2vec.search_architecture(
    df=df, base_config=base_config, search_config=search_config
)

assert "n_layers" in best_architecture
assert "hidden_units" in best_architecture

Searching costs one training run per trial — budget max_trials accordingly.


Saving and reusing a model

An embedding is only reproducible if the preprocessing is reproducible too. A saved Row2Vec model carries its encoders, imputers, and scalers with it, so embedding new rows later stays consistent with training.

import tempfile
from pathlib import Path

import row2vec

df = row2vec.generate_synthetic_data(100)
base = Path(tempfile.mkdtemp()) / "model"

embeddings, saved = row2vec.train_and_save_model(
    df, base_path=str(base), mode="pca", embedding_dim=2
)

model = row2vec.load_model(saved)
new_rows = row2vec.generate_synthetic_data(20, seed=99)

info = row2vec.inspect_model(saved)  # reads only the manifest

assert embeddings.shape == (100, 2)
assert model.predict(new_rows).shape == (20, 2)
assert info["mode"] == "pca" and info["format"] == "row2vec-model"

A saved model is a single .r2v file (a zip archive). It holds the fitted preprocessing, the projector or trained encoder, the configuration, and a manifest.json with the training metadata and the library versions that wrote it. inspect_model reads that manifest without loading anything else, and unzip -l model.r2v shows the members.

Loading runs no code from the file. There is no loader script and no pickle: scikit-learn objects are read with skops, which refuses any type outside a short allow-list kept in row2vec, and the encoder of a neural mode is read in Keras's own format with safe_mode and only Dense and Dropout layers permitted. A file that names anything else raises ModelFormatError instead of loading. This defends against a malicious file, not an unauthenticated one: a checksum for each member catches corruption, but anyone who can rewrite the file can rewrite the checksums too, so still load models only from sources you trust. See SECURITY.md.

Two things to know. A model is read back by the same major row2vec and scikit-learn that wrote it; the manifest records both, and a file from a newer format version is refused with a request to upgrade. And a saved model contains the category values and summary statistics the preprocessing learned from your data (not the rows themselves), so treat it as derived from the data when you share it.


Integrations

scikit-learn pipelines

Row2VecTransformer is a standard transformer: put it in a Pipeline and it behaves like any other step.

from sklearn.pipeline import Pipeline

import row2vec
from row2vec import EmbeddingConfig, Row2VecTransformer

df = row2vec.generate_synthetic_data(100)

pipeline = Pipeline(
    [("embed", Row2VecTransformer(config=EmbeddingConfig(mode="pca", embedding_dim=2)))]
)
transformed = pipeline.fit_transform(df)

assert transformed.shape == (100, 2)

The pandas accessor

Importing row2vec registers a .row2vec accessor on DataFrame, which is convenient in a notebook.

import row2vec  # importing registers the accessor

df = row2vec.generate_synthetic_data(80)
embeddings = df.row2vec.pca(dim=2)

assert embeddings.shape == (80, 2)

The command line

For batch work there is no need to write Python at all:

# Embed a file in one step
row2vec annotate --input data.csv --output embeddings.csv --mode pca --dim 5

# Train a reusable model, then apply it to new data
row2vec train --input data.csv --output model.py --mode unsupervised --dim 10
row2vec predict --input new.csv --model model.py --output predictions.csv

Practical notes

Reproducibility

Every mode takes a seed. The same seed and the same input give the same embedding.

import row2vec

df = row2vec.generate_synthetic_data(80)

first = row2vec.learn_embedding(df, mode="pca", embedding_dim=2, seed=1305)
second = row2vec.learn_embedding(df, mode="pca", embedding_dim=2, seed=1305)

assert first.equals(second)

Neural modes are seeded the same way, though exact floating-point results can still differ across platforms and library versions.

Configuration objects

For anything beyond a few arguments, EmbeddingConfig groups the settings so they can be reused, serialised, and version-controlled.

import row2vec
from row2vec import EmbeddingConfig, NeuralConfig

config = EmbeddingConfig(
    mode="unsupervised",
    embedding_dim=4,
    neural=NeuralConfig(hidden_units=[32, 16], max_epochs=2, dropout_rate=0.1),
)

df = row2vec.generate_synthetic_data(100)
embeddings = row2vec.learn_embedding_v2(df, config)

assert embeddings.shape == (100, 4)

Logging

Training emits structured logs. Turn them off for quiet runs, or point them at a file with log_file=.

import row2vec

df = row2vec.generate_synthetic_data(50)
embeddings = row2vec.learn_embedding(df, mode="pca", embedding_dim=2, enable_logging=False)

assert embeddings.shape == (50, 2)

What is not supported

Columns of any dtype other than numeric, categorical, boolean, datetime or declared text (periods, intervals) are ignored, with a warning that names them.


Next steps

  • API Reference — every public class and function, generated from the source.
  • Home — the one-minute overview and the method table.