Skip to content

API Reference

This page is generated automatically from the docstrings in the row2vec package, so it always matches the installed code.

Row2Vec: A library for learning embeddings from tabular data.

This library provides both neural network and classical machine learning approaches for creating vector embeddings from tabular datasets.

ArchitectureSearchConfig dataclass

ArchitectureSearchConfig(method: str = 'random', max_trials: int = 30, max_time: float | None = 1800, patience: int = 10, min_improvement: float = 0.01, layer_range: tuple[int, int] = (1, 4), max_layers: int = 4, width_options: list[int] = (lambda: [32, 64, 128, 256, 512])(), dropout_options: list[float] = (lambda: [0.0, 0.1, 0.2, 0.3, 0.4, 0.5])(), activation_options: list[str] = (lambda: ['relu', 'elu', 'swish'])(), initial_epochs: int = 10, intermediate_epochs: int = 25, final_epochs: int = 50, top_k_intermediate: int = 10, top_k_final: int = 3, reconstruction_weight: float = 0.4, clustering_weight: float = 0.3, efficiency_weight: float = 0.2, stability_weight: float = 0.1, verbose: bool = True, random_seed: int | None = None, return_full_history: bool = False)

Configuration for automatic neural architecture search.

This class defines the search space, evaluation criteria, and stopping conditions for finding optimal neural network architectures.

ArchitectureSearcher

ArchitectureSearcher(config: ArchitectureSearchConfig)

Main class for performing neural architecture search.

Implements multiple search strategies to find optimal neural network architectures for embedding generation tasks.

search

search(df: DataFrame, base_config: EmbeddingConfig, target_column: str | None = None) -> ArchitectureSearchResult

Perform architecture search on the given dataset.

Parameters:

Name Type Description Default
df DataFrame

Input dataframe for embedding generation

required
base_config EmbeddingConfig

Base embedding configuration

required
target_column str | None

Optional target column for supervised evaluation

None

Returns:

Type Description
ArchitectureSearchResult

ArchitectureSearchResult containing the best architecture and metadata

ArchitectureSearchResult

ArchitectureSearchResult(best_architecture: dict[str, Any], best_score: float, search_history: list[dict[str, Any]], total_time: float, trials_completed: int)

Container for architecture search results.

summary

summary() -> dict[str, Any]

Get a summary of the search results.

AutoDimensionSelector

AutoDimensionSelector(methods: list[str] | None = None, performance_weight: float = 0.4, efficiency_weight: float = 0.3, intrinsic_weight: float = 0.3, max_dimension: int | None = None, min_dimension: int = 2, n_trials: int = 5, verbose: bool = True, random_state: int = 1305)

Automatically selects optimal embedding dimensions using multiple strategies.

Combines data-driven analysis, performance optimization, and heuristic rules to determine the best embedding dimension for a given dataset.

Initialize automatic dimension selector.

Parameters:

Name Type Description Default
methods list[str] | None

List of selection methods to use

None
performance_weight float

Weight for performance-based selection

0.4
efficiency_weight float

Weight for efficiency considerations

0.3
intrinsic_weight float

Weight for intrinsic dimensionality estimation

0.3
max_dimension int | None

Maximum dimension to consider (auto if None)

None
min_dimension int

Minimum dimension to consider

2
n_trials int

Number of trials for performance evaluation

5
random_state int

Seed for every estimator this selector fits

1305
verbose bool

Whether to show selection progress

True

select_dimension

select_dimension(df: DataFrame, config: EmbeddingConfig, target_column: str | None = None, candidate_dims: list[int] | None = None) -> tuple[int, dict[str, Any]]

Select optimal embedding dimension for the given data.

Parameters:

Name Type Description Default
df DataFrame

Input dataframe

required
config EmbeddingConfig

Base embedding configuration (dimension will be overridden)

required
target_column str | None

Optional target for supervised evaluation

None
candidate_dims list[int] | None

Specific dimensions to evaluate (auto-generated if None)

None

Returns:

Type Description
tuple[int, dict[str, Any]]

Tuple of (optimal_dimension, selection_metadata)

CategoricalAnalyzer

CategoricalAnalyzer(config: CategoricalEncodingConfig)

Analyzes categorical data to recommend optimal encoding strategies.

analyze_column

analyze_column(series: Series, target: Series | None = None) -> dict[str, Any]

Analyze a categorical column to recommend encoding strategy.

Parameters:

Name Type Description Default
series Series

Categorical column to analyze

required
target Series

Target variable for correlation analysis

None

Returns:

Type Description
dict[str, Any]

Dict[str, Any] Analysis results and strategy recommendation

CategoricalEncoder

CategoricalEncoder(config: CategoricalEncodingConfig | None = None)

Bases: BaseEstimator, TransformerMixin

Intelligent categorical encoder with adaptive strategy selection.

This encoder analyzes categorical data characteristics and automatically selects optimal encoding strategies while providing full control for advanced users.

fit

fit(X: DataFrame, y: Series | None = None) -> CategoricalEncoder

Fit the categorical encoder on training data.

Parameters:

Name Type Description Default
X DataFrame

Categorical features to encode

required
y Series

Target variable for supervised encoding strategies

None

Returns:

Name Type Description
self CategoricalEncoder

Fitted encoder instance

fit_transform

fit_transform(X: DataFrame, y: Series | None = None, **fit_params: Any) -> pd.DataFrame

Fit on X and encode it without leaking the target.

TransformerMixin would give us fit(X, y).transform(X), and transform deliberately uses the full-data category means - correct for rows the encoder has never seen, badly leaky for the rows it was just fitted on. Target-encoded columns therefore take their cross-fitted values here, matching :class:sklearn.preprocessing.TargetEncoder.

Parameters:

Name Type Description Default
X DataFrame

Categorical features to encode.

required
y Series

Target for supervised strategies.

None
**fit_params Any

Ignored; present for the sklearn signature.

{}

Returns:

Type Description
DataFrame

pd.DataFrame: Encoded training features.

transform

transform(X: DataFrame) -> pd.DataFrame

Transform categorical data using fitted encoders.

Parameters:

Name Type Description Default
X DataFrame

Categorical data to transform

required

Returns:

Type Description
DataFrame

pd.DataFrame Encoded categorical data

get_feature_names_out

get_feature_names_out(input_features: list[str] | None = None) -> list[str]

Get output feature names for transformation.

get_analysis_report

get_analysis_report() -> dict[str, dict[str, Any]]

Get detailed analysis report for all columns.

CategoricalEncodingConfig dataclass

CategoricalEncodingConfig(encoding_strategy: str = 'adaptive', onehot_threshold: int = 20, target_threshold: int = 100, entity_threshold: int = 1000, correlation_threshold: float = 0.1, drop_identifiers: bool = True, min_identifier_rows: int = 50, target_smoothing: float = 1.0, target_noise: float = 0.01, target_cv_folds: int = 5, embedding_dim_ratio: float = 0.5, min_embedding_dim: int = 2, max_embedding_dim: int = 50, entity_epochs: int = 50, entity_batch_size: int = 256, prefer_speed: bool = True, preserve_interpretability: bool = False, enable_feature_selection: bool = False, feature_importance_threshold: float = 0.01, custom_strategies: dict[str, str] = dict(), handle_unknown: str = 'ignore', random_state: int = 42)

Configuration for intelligent categorical encoding strategies.

This class provides comprehensive control over how categorical variables are encoded, with intelligent defaults that automatically select optimal strategies based on data characteristics while allowing expert users to fine-tune every aspect.

encoding_strategy class-attribute instance-attribute

encoding_strategy: str = 'adaptive'

Encoding strategy selection. Options: - "adaptive": Automatically selects best strategy based on data analysis - "onehot": One-hot encoding for all categorical features - "target": Target encoding for all categorical features - "entity": Entity embeddings for all categorical features - "ordinal": Ordinal encoding (assumes natural order) - "mixed": Use custom strategies per column (requires custom_strategies)

onehot_threshold class-attribute instance-attribute

onehot_threshold: int = 20

Use OneHot encoding if cardinality <= this threshold and correlation is low.

target_threshold class-attribute instance-attribute

target_threshold: int = 100

Use target encoding if cardinality is between onehot_threshold and this value.

entity_threshold class-attribute instance-attribute

entity_threshold: int = 1000

Use entity embeddings if cardinality > target_threshold and <= this value.

correlation_threshold class-attribute instance-attribute

correlation_threshold: float = 0.1

Minimum mutual information score to prefer target/entity over onehot.

drop_identifiers class-attribute instance-attribute

drop_identifiers: bool = True

Drop categorical columns whose every value is distinct (names, ids).

Such a column cannot generalise: no value seen at fit time ever recurs, so whatever is learned about it describes only the training rows, and every new row arrives with an unseen value. Encoding it anyway adds noise dimensions. Applies from min_identifier_rows rows up.

min_identifier_rows class-attribute instance-attribute

min_identifier_rows: int = 50

Fewest rows at which an all-distinct column is judged an identifier.

target_smoothing class-attribute instance-attribute

target_smoothing: float = 1.0

Bayesian smoothing factor for target encoding. Higher values = more smoothing.

target_noise class-attribute instance-attribute

target_noise: float = 0.01

Gaussian noise standard deviation added to target encodings to prevent overfitting.

target_cv_folds class-attribute instance-attribute

target_cv_folds: int = 5

Number of cross-validation folds for target encoding to prevent data leakage.

embedding_dim_ratio class-attribute instance-attribute

embedding_dim_ratio: float = 0.5

Embedding dimension as ratio of sqrt(cardinality). Controls embedding size.

min_embedding_dim class-attribute instance-attribute

min_embedding_dim: int = 2

Minimum embedding dimension for entity embeddings.

max_embedding_dim class-attribute instance-attribute

max_embedding_dim: int = 50

Maximum embedding dimension for entity embeddings.

entity_epochs class-attribute instance-attribute

entity_epochs: int = 50

Number of training epochs for entity embedding networks.

entity_batch_size class-attribute instance-attribute

entity_batch_size: int = 256

Batch size for entity embedding training.

prefer_speed class-attribute instance-attribute

prefer_speed: bool = True

Whether to prefer faster methods over more accurate but slower ones.

preserve_interpretability class-attribute instance-attribute

preserve_interpretability: bool = False

Whether to prefer interpretable encodings (OneHot/Ordinal) when possible.

enable_feature_selection class-attribute instance-attribute

enable_feature_selection: bool = False

Whether to enable automatic feature selection based on importance.

feature_importance_threshold class-attribute instance-attribute

feature_importance_threshold: float = 0.01

Minimum feature importance score to keep feature (only if enable_feature_selection=True).

custom_strategies class-attribute instance-attribute

custom_strategies: dict[str, str] = field(default_factory=dict)

Custom encoding strategy for specific columns. Format: {column_name: strategy}

handle_unknown class-attribute instance-attribute

handle_unknown: str = 'ignore'

How to handle unknown categories. Options: 'ignore', 'error', 'infrequent_if_exist'

random_state class-attribute instance-attribute

random_state: int = 42

Random state for reproducible results.

EntityEmbeddingTrainer

EntityEmbeddingTrainer(config: CategoricalEncodingConfig)

Trains entity embeddings for high-cardinality categorical features.

fit_column_embedding

fit_column_embedding(series: Series, target: Series | None = None, embedding_dim: int = 10) -> NDArray[Any]

Train entity embeddings for a categorical column.

Parameters:

Name Type Description Default
series Series

Categorical column to embed

required
target Series

Target variable for supervised embedding

None
embedding_dim int

Dimension of embedding vectors

10

Returns:

Type Description
NDArray[Any]

np.ndarray Trained embedding matrix of shape (cardinality, embedding_dim)

TargetEncoder

TargetEncoder(config: CategoricalEncodingConfig)

Implements Bayesian target encoding with cross-validation.

fit_transform

fit_transform(series: Series, target: Series) -> pd.Series

Fit target encoder and transform the series.

Parameters:

Name Type Description Default
series Series

Categorical column to encode

required
target Series

Target variable

required

Returns:

Type Description
Series

pd.Series Target-encoded values

transform

transform(series: Series) -> pd.Series

Transform new data using fitted encodings.

ClassicalConfig dataclass

ClassicalConfig(n_neighbors: int = 15, min_dist: float = 0.1, perplexity: float = 30.0, n_iter: int = 1000)

Configuration for classical ML dimensionality reduction methods.

ContrastiveConfig dataclass

ContrastiveConfig(loss_type: str = 'triplet', similar_pairs: list[tuple[int, int]] | None = None, dissimilar_pairs: list[tuple[int, int]] | None = None, auto_pairs: str | None = None, margin: float = 1.0, negative_samples: int = 5)

Configuration for contrastive learning.

EmbeddingConfig dataclass

EmbeddingConfig(embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, seed: int = 1305, verbose: bool = False, neural: NeuralConfig = NeuralConfig(), classical: ClassicalConfig = ClassicalConfig(), contrastive: ContrastiveConfig = ContrastiveConfig(), scaling: ScalingConfig = ScalingConfig(), logging: LoggingConfig = LoggingConfig(), preprocessing: PreprocessingConfig = PreprocessingConfig())

Complete configuration for embedding learning.

from_dict classmethod

from_dict(config_dict: dict[str, Any]) -> EmbeddingConfig

Create config from dictionary (e.g., from YAML).

The caller's dictionary is left untouched; an earlier version popped keys straight out of it, so reusing a config dict silently produced a different config the second time.

from_yaml classmethod

from_yaml(yaml_path: str | Path) -> EmbeddingConfig

Create config from YAML file.

to_dict

to_dict() -> dict[str, Any]

Convert config to dictionary.

to_yaml

to_yaml(yaml_path: str | Path) -> None

Save config to YAML file.

LoggingConfig dataclass

LoggingConfig(level: str = 'INFO', file: str | None = None, enabled: bool = True)

Configuration for logging and output.

NeuralConfig dataclass

NeuralConfig(max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, activation: str = 'relu', early_stopping: bool = True)

Configuration for neural network-based embedding methods.

PreprocessingConfig dataclass

PreprocessingConfig(handle_missing: str = 'auto', numeric_scaling: str = 'standard', categorical_encoding_strategy: str = 'adaptive', categorical_onehot_threshold: int = 20, categorical_target_threshold: int = 100, categorical_entity_threshold: int = 1000, text_columns: list[str] = list(), text_dim: int = 16, text_encoder: Callable[[list[str]], Any] | None = None)

Configuration for data preprocessing including categorical encoding.

ScalingConfig dataclass

ScalingConfig(method: str | None = None, range: tuple[float, float] | None = None)

Configuration for embedding scaling/normalization.

AdaptiveImputer

AdaptiveImputer(config: ImputationConfig)

Bases: BaseEstimator

Adaptive imputer that automatically selects and applies appropriate imputation strategies based on data characteristics.

fit

fit(X: DataFrame, y: Any = None) -> AdaptiveImputer

Fit the adaptive imputer to the data.

Parameters:

Name Type Description Default
X DataFrame

Input DataFrame with potential missing values

required
y Any

Ignored, present for API compatibility

None

Returns:

Name Type Description
self AdaptiveImputer

Fitted imputer

transform

transform(X: DataFrame) -> pd.DataFrame

Transform the data by applying imputation strategies.

Parameters:

Name Type Description Default
X DataFrame

Input DataFrame with potential missing values

required

Returns:

Type Description
DataFrame

DataFrame with missing values imputed

fit_transform

fit_transform(X: DataFrame, y: Any = None, **fit_params: Any) -> pd.DataFrame

Fit the imputer and transform the data in one step.

get_imputation_report

get_imputation_report() -> dict[str, Any]

Get detailed report about the imputation process.

Returns:

Type Description
dict[str, Any]

Dict containing analysis and imputation details

ImputationConfig dataclass

ImputationConfig(numeric_strategy: str = 'adaptive', categorical_strategy: str = 'adaptive', prefer_speed: bool = True, missing_threshold: float = 0.7, row_missing_threshold: float = 0.9, knn_neighbors: int = 5, preserve_missing_patterns: bool = False, missing_indicator_suffix: str = '_was_missing', auto_detect_patterns: bool = True, warn_high_missingness: bool = True, categorical_fill_value: str = 'Missing')

Configuration for intelligent missing value imputation strategies.

This class provides comprehensive control over how missing values are handled, with sensible defaults that work well for most datasets while allowing power users to fine-tune every aspect of the imputation process.

numeric_strategy class-attribute instance-attribute

numeric_strategy: str = 'adaptive'

Numeric imputation strategy. Options: - "adaptive": Automatically selects best strategy based on missing percentage - "mean": Mean imputation (fastest, good for <10% missing) - "median": Median imputation (robust to outliers, good for 10-30% missing) - "knn": K-nearest neighbors imputation (better for >30% missing) - "iterative": MICE-style iterative imputation (best quality, slowest)

categorical_strategy class-attribute instance-attribute

categorical_strategy: str = 'adaptive'

Categorical imputation strategy. Options: - "adaptive": Automatically selects best strategy based on data characteristics - "mode": Most frequent value imputation - "constant": Fill with specified constant value - "missing_category": Create explicit "Missing" category

prefer_speed class-attribute instance-attribute

prefer_speed: bool = True

Whether to prefer faster methods over more accurate but slower ones. When True, uses simpler strategies by default. When False, prefers more sophisticated methods even if they take longer.

missing_threshold class-attribute instance-attribute

missing_threshold: float = 0.7

Columns with more than this fraction of missing values will be flagged. Conservative default of 0.7 to avoid dropping useful but sparse columns.

row_missing_threshold class-attribute instance-attribute

row_missing_threshold: float = 0.9

Rows with more than this fraction of missing values will be flagged. Very conservative default to avoid losing data.

knn_neighbors class-attribute instance-attribute

knn_neighbors: int = 5

Number of neighbors for KNN imputation. Should be odd to avoid ties.

preserve_missing_patterns class-attribute instance-attribute

preserve_missing_patterns: bool = False

Whether to preserve missing patterns when they might be informative.

When True, adds binary indicator columns for originally missing values. This is useful when missingness itself carries information (e.g., customers not providing income information might be systematically different).

Example

Original: [1.0, NaN, 3.0] -> After imputation: [1.0, 2.0, 3.0] With preservation: adds column [False, True, False] indicating missingness

missing_indicator_suffix class-attribute instance-attribute

missing_indicator_suffix: str = '_was_missing'

Suffix for missing indicator columns when preserve_missing_patterns=True.

auto_detect_patterns class-attribute instance-attribute

auto_detect_patterns: bool = True

Whether to automatically analyze missing data patterns and adjust strategies.

warn_high_missingness class-attribute instance-attribute

warn_high_missingness: bool = True

Whether to warn users about columns/rows with high missing percentages.

categorical_fill_value class-attribute instance-attribute

categorical_fill_value: str = 'Missing'

Fill value when using 'constant' strategy for categorical data.

MissingPatternAnalyzer

MissingPatternAnalyzer(config: ImputationConfig)

Analyzes missing data patterns to inform imputation strategy selection.

analyze

analyze(df: DataFrame) -> dict[str, Any]

Analyze missing data patterns in the DataFrame.

Parameters:

Name Type Description Default
df DataFrame

Input DataFrame to analyze

required

Returns:

Type Description
dict[str, Any]

Dict containing analysis results and recommendations

Row2VecLogger

Row2VecLogger(name: str = 'row2vec', level: str = 'INFO', log_file: str | Path | None = None, include_performance: bool = True, include_memory: bool = True)

Centralized logging system for Row2Vec operations.

Provides structured logging for training progress, debug information, and performance metrics with configurable output formats and levels.

Initialize Row2Vec logger.

Parameters:

Name Type Description Default
name str

Logger name

'row2vec'
level str

Logging level (DEBUG, INFO, WARNING, ERROR)

'INFO'
log_file str | Path | None

Optional file path for logging output

None
include_performance bool

Whether to include performance metrics

True
include_memory bool

Whether to include memory usage tracking

True

close

close() -> None

Close and detach the log files this logger opened.

Call this when you are finished with a file-backed logger. Until you do, the file stays open: on Windows that means it cannot be deleted or renamed, and on every platform the handle is held for the life of the process.

Safe to call more than once.

Examples:

>>> import tempfile
>>> from pathlib import Path
>>> from row2vec import get_logger
>>> path = Path(tempfile.mkdtemp()) / "run.log"
>>> logger = get_logger(name="row2vec.example", log_file=path)
>>> logger.close()
>>> path.exists()
True

start_training

start_training(**kwargs: Any) -> None

Log training start with configuration details.

start_epoch

start_epoch(epoch: int, total_epochs: int) -> None

Log epoch start.

log_epoch_metrics

log_epoch_metrics(epoch: int, loss: float, val_loss: float | None = None, additional_metrics: dict[str, float] | None = None) -> None

Log epoch completion with metrics.

log_early_stopping

log_early_stopping(epoch: int, reason: str) -> None

Log early stopping event.

end_training

end_training(final_loss: float, total_epochs: int) -> None

Log training completion with summary.

log_data_preprocessing

log_data_preprocessing(df_shape: tuple[int, int], processing_steps: list[str]) -> None

Log data preprocessing information.

log_preprocessing_result

log_preprocessing_result(original_shape: tuple[int, int], processed_shape: tuple[int, int], processing_time: float) -> None

Log preprocessing completion.

log_model_architecture

log_model_architecture(model_summary: str) -> None

Log model architecture details.

log_embedding_stats

log_embedding_stats(embeddings: DataFrame) -> None

Log embedding statistics.

log_performance_warning

log_performance_warning(message: str) -> None

Log performance-related warnings.

log_validation_issue

log_validation_issue(message: str) -> None

Log validation or data quality issues.

log_debug_info

log_debug_info(message: str, data: dict[str, Any] | None = None) -> None

Log debug information with optional data context.

log_error

log_error(error: Exception, context: str | None = None) -> None

Log error with context information.

log_completion

log_completion(message: str = 'Embedding generation completed successfully!') -> None

Log completion of embedding generation.

PipelineBuilder

PipelineBuilder(config: EmbeddingConfig | None = None)

Intelligent pipeline builder that analyzes data and constructs optimal preprocessing pipelines with adaptive strategies.

build_preprocessing_pipeline

build_preprocessing_pipeline(df: DataFrame, target: Series | None = None, mode: str = 'unsupervised') -> tuple[ColumnTransformer, dict[str, Any]]

Build intelligent preprocessing pipeline based on data analysis.

Parameters:

Name Type Description Default
df DataFrame

Input dataset to analyze

required
target Series

Target variable for supervised preprocessing

None
mode str

Embedding mode that influences preprocessing strategy

'unsupervised'

Returns:

Type Description
tuple[ColumnTransformer, dict[str, Any]]

Tuple[ColumnTransformer, Dict[str, Any]] Fitted preprocessing pipeline and analysis report

get_analysis_report

get_analysis_report() -> dict[str, Any]

Get detailed analysis report of the dataset.

get_pipeline_description

get_pipeline_description() -> dict[str, Any]

Get human-readable description of the constructed pipeline.

ModelFormatError

Bases: ValueError

A saved model is not in a form that can be loaded safely.

Row2VecModel

Row2VecModel(config: EmbeddingConfig | None = None, *, aggregate_by_reference: bool = False)

A fitted row2vec embedder.

Holds every part needed to embed new rows the same way the training rows were embedded: the fitted preprocessing pipeline, the fitted projector, the encoder view for neural modes, and the fitted embedding scaler.

Parameters:

Name Type Description Default
config EmbeddingConfig

Configuration to fit with. Defaults to :class:EmbeddingConfig.

None
aggregate_by_reference bool

mode="target" only. When True, fit_transform returns one row per distinct reference value instead of one row per input row. Off by default, so target mode carries df's index like every other mode.

False

Attributes:

Name Type Description
preprocessor_ ColumnTransformer

The fitted preprocessing pipeline.

projector_ object

The fitted PCA/UMAP/t-SNE estimator, or the trained Keras model.

encoder_ object or None

For neural modes, the sub-model mapping processed features to the embedding. None for classical modes.

embedding_scaler_ EmbeddingScaler

The fitted embedding scaler.

feature_names_in_ list[Hashable]

Column labels seen during fit.

metadata object or None

Populated by :mod:row2vec.serialization on save/load.

Examples:

>>> import row2vec
>>> df = row2vec.generate_synthetic_data(60)
>>> model = Row2VecModel(row2vec.EmbeddingConfig(embedding_dim=2, mode="pca"))
>>> train = model.fit_transform(df)
>>> list(train.index) == list(df.index)
True

transform reuses the fitted basis rather than refitting:

>>> model.transform(df.head(1)).shape
(1, 2)

mode property

mode: str

The embedding mode this model was configured with.

spec property

spec: ModeSpec

The :class:ModeSpec for this model's mode.

is_fitted property

is_fitted: bool

Whether :meth:fit has run.

fit

fit(df: DataFrame) -> Row2VecModel

Fit the preprocessor, projector, and embedding scaler on df.

Parameters:

Name Type Description Default
df DataFrame

Training data.

required

Returns:

Name Type Description
Row2VecModel Row2VecModel

self, so calls can be chained.

fit_transform

fit_transform(df: DataFrame) -> pd.DataFrame

Fit on df and return its embeddings in one training pass.

Parameters:

Name Type Description Default
df DataFrame

Training data.

required

Returns:

Type Description
DataFrame

pd.DataFrame: Embeddings indexed by df.index, unless

DataFrame

aggregate_by_reference is set for mode="target".

transform

transform(df: DataFrame) -> pd.DataFrame

Embed rows using the state learned during :meth:fit.

Parameters:

Name Type Description Default
df DataFrame

Rows to embed. May be a single row.

required

Returns:

Type Description
DataFrame

pd.DataFrame: Embeddings indexed by df.index.

Raises:

Type Description
RuntimeError

If the model has not been fitted.

NotImplementedError

If the mode has no out-of-sample extension.

predict

predict(df: DataFrame, validate_schema: bool = True) -> pd.DataFrame

Embed df with a saved model.

Parameters:

Name Type Description Default
df DataFrame

Rows to embed.

required
validate_schema bool

Whether to check df against the schema recorded at training time. Ignored when the model carries no metadata, which is the case before it has been saved.

True

Returns:

Type Description
DataFrame

pd.DataFrame: Embeddings indexed by df.index, identical to what

DataFrame

training returned for the same rows.

validate_input_schema

validate_input_schema(df: DataFrame, strict: bool = True) -> bool

Check df against the schema recorded when the model was saved.

Parameters:

Name Type Description Default
df DataFrame

Input to validate.

required
strict bool

Raise on mismatch rather than returning False.

True

Returns:

Name Type Description
bool bool

Whether the schema matches.

Raises:

Type Description
ValueError

If strict and the schema does not match, or no schema was recorded.

to_state

to_state() -> dict[str, Any]

Everything needed to rebuild this model, for pickling.

Returns:

Type Description
dict[str, Any]

dict[str, Any]: The fitted state, keyed by attribute name.

from_state classmethod

from_state(state: dict[str, Any]) -> Row2VecModel

Rebuild a fitted model from :meth:to_state output.

Parameters:

Name Type Description Default
state dict[str, Any]

A previously saved state dictionary.

required

Returns:

Name Type Description
Row2VecModel Row2VecModel

The restored model, ready to transform.

Row2VecModelMetadata

Row2VecModelMetadata(embedding_dim: int, mode: str, reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, early_stopping: bool = True, seed: int = 1305, scale_method: str | None = None, scale_range: tuple[float, float] | None = None, n_neighbors: int = 15, perplexity: float = 30.0, min_dist: float = 0.1, n_iter: int = 1000, training_history: dict[str, Any] | None = None, final_loss: float | None = None, epochs_trained: int | None = None, training_time: float | None = None, original_columns: list[str] | None = None, preprocessed_feature_names: list[str] | None = None, data_shape: tuple[int, int] | None = None, data_types: dict[str, str] | None = None, expected_schema: dict[str, Any] | None = None)

Container for Row2Vec model training metadata.

to_dict

to_dict() -> dict[str, Any]

Convert metadata to dictionary for serialization.

from_dict classmethod

from_dict(data: dict[str, Any]) -> Row2VecModelMetadata

Create metadata from dictionary.

Row2VecClassifier

Row2VecClassifier(embedding_dim: int = 10, mode: str = 'pca', reference_column: str | None = None, classifier: Any = None, seed: int = 1305, embedding_config: EmbeddingConfig | None = None)

Bases: ClassifierMixin, BaseEstimator

Classify rows using Row2Vec embeddings as features.

Parameters:

Name Type Description Default
embedding_dim int

Width of the embedding space.

10
mode str

Embedding mode. See :class:Row2VecTransformer.

'pca'
reference_column str

Label column for mode="target".

None
classifier sklearn estimator

Downstream classifier. Cloned before fitting, so the caller's instance is left alone. Defaults to LogisticRegression.

None
seed int

Random seed.

1305
embedding_config EmbeddingConfig

Complete embedding configuration; overrides the individual arguments.

None

Attributes:

Name Type Description
transformer_ Row2VecTransformer

The fitted embedder.

classifier_ sklearn estimator

The fitted downstream classifier.

classes_ ndarray

Class labels seen during fit.

Examples:

>>> import row2vec
>>> from row2vec.sklearn import Row2VecClassifier
>>> df = row2vec.generate_synthetic_data(80)
>>> y = df["Country"]
>>> clf = Row2VecClassifier(embedding_dim=2, mode="pca")
>>> _ = clf.fit(df.drop(columns=["Country"]), y)
>>> len(clf.predict(df.drop(columns=["Country"])))
80

fit

fit(X: Any, y: Any) -> Row2VecClassifier

Fit the embedder, then the downstream classifier on its output.

Parameters:

Name Type Description Default
X Any

Training data.

required
y Any

Target labels.

required

Returns:

Name Type Description
Row2VecClassifier Row2VecClassifier

self.

predict

predict(X: Any) -> np.ndarray[Any, Any]

Predict labels for X.

Parameters:

Name Type Description Default
X Any

Data to classify.

required

Returns:

Type Description
ndarray[Any, Any]

ndarray of predicted labels.

predict_proba

predict_proba(X: Any) -> np.ndarray[Any, Any]

Predict class probabilities for X.

Parameters:

Name Type Description Default
X Any

Data to classify.

required

Returns:

Type Description
ndarray[Any, Any]

ndarray of shape (n_samples, n_classes).

Row2VecTransformer

Row2VecTransformer(embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, activation: str = 'relu', seed: int = 1305, config: EmbeddingConfig | None = None)

Bases: TransformerMixin, BaseEstimator

Scikit-learn transformer producing Row2Vec embeddings.

TransformerMixin comes first in the bases deliberately: scikit-learn reads the transformer tags off the MRO, and with BaseEstimator first check_estimator refuses to run at all.

Parameters:

Name Type Description Default
embedding_dim int

Width of the embedding space.

10
mode str

"unsupervised", "target", "pca", "umap" or "contrastive". "tsne" is rejected at fit time because it cannot transform unseen rows, and a transformer that cannot transform has no place in a pipeline.

'unsupervised'
reference_column str

Label column for mode="target".

None
max_epochs int

Training epoch ceiling for neural modes.

50
batch_size int

Training batch size.

64
dropout_rate float

Dropout after each hidden layer.

0.2
hidden_units int | list[int]

Hidden layer width, or widths.

128
activation str

Hidden-layer activation.

'relu'
seed int

Random seed.

1305
config EmbeddingConfig

A complete configuration. When given, the other arguments are ignored.

None

Attributes:

Name Type Description
model_ Row2VecModel

The fitted model. transform projects with it rather than refitting.

config_ EmbeddingConfig

The configuration actually used.

n_features_in_ int

Number of columns seen during fit.

feature_names_in_ ndarray

Column names seen during fit.

Examples:

>>> import row2vec
>>> from row2vec.sklearn import Row2VecTransformer
>>> df = row2vec.generate_synthetic_data(60)
>>> transformer = Row2VecTransformer(embedding_dim=2, mode="pca")
>>> transformer.fit_transform(df).shape
(60, 2)

A fitted transformer embeds a single unseen row:

>>> transformer.transform(df.head(1)).shape
(1, 2)

fit

fit(X: Any, y: Any = None) -> Row2VecTransformer

Fit the embedding model.

Parameters:

Name Type Description Default
X Any

Training data.

required
y Any

Ignored; present for the sklearn signature.

None

Returns:

Name Type Description
Row2VecTransformer Row2VecTransformer

self.

Raises:

Type Description
ValueError

If mode cannot support transform.

transform

transform(X: Any) -> np.ndarray[Any, Any]

Project data into the fitted embedding space.

Parameters:

Name Type Description Default
X Any

Data to transform. A single row is fine.

required

Returns:

Type Description
ndarray[Any, Any]

ndarray of shape (n_samples, embedding_dim).

fit_transform

fit_transform(X: Any, y: Any = None, **fit_params: Any) -> np.ndarray[Any, Any]

Fit and transform in a single training pass.

Parameters:

Name Type Description Default
X Any

Training data.

required
y Any

Ignored; present for the sklearn signature.

None
**fit_params Any

Ignored; present for the sklearn signature.

{}

Returns:

Type Description
ndarray[Any, Any]

ndarray of shape (n_samples, embedding_dim).

get_feature_names_out

get_feature_names_out(input_features: ndarray[Any, Any] | None = None) -> np.ndarray[Any, Any]

Names of the embedding columns.

Parameters:

Name Type Description Default
input_features ndarray[Any, Any] | None

Ignored; present for the sklearn signature.

None

Returns:

Type Description
ndarray[Any, Any]

ndarray of str: row2vec_0 through row2vec_{n-1}.

learn_embedding_classical

learn_embedding_classical(df: DataFrame, method: str = 'pca', embedding_dim: int = 10, **overrides: Any) -> pd.DataFrame

Learn classical ML embeddings (PCA, t-SNE, UMAP) with optimized defaults.

learn_embedding_contrastive

learn_embedding_contrastive(df: DataFrame, **overrides: Any) -> pd.DataFrame

Convenience function for contrastive learning with optimized defaults.

learn_embedding_target

learn_embedding_target(df: DataFrame, reference_column: str, embedding_dim: int = 10, **overrides: Any) -> pd.DataFrame

Learn target-based embeddings with optimized defaults.

learn_embedding_unsupervised

learn_embedding_unsupervised(df: DataFrame, embedding_dim: int = 10, **overrides: Any) -> pd.DataFrame

Learn unsupervised embeddings with optimized defaults.

learn_embedding_v2

learn_embedding_v2(df: DataFrame, config: EmbeddingConfig | None = None, auto_architecture: bool = False, architecture_search_config: Optional[ArchitectureSearchConfig] = None, **config_overrides: Any) -> pd.DataFrame

Modern config-based API for learning embeddings from tabular data.

This is the new recommended API that uses configuration objects instead of long parameter lists. It provides better organization, type safety, and extensibility.

Parameters:

Name Type Description Default
df DataFrame

Input DataFrame containing the data to embed

required
config EmbeddingConfig | None

Complete embedding configuration. If None, default config is used.

None
auto_architecture bool

Enable automatic neural architecture search for neural modes

False
architecture_search_config Optional[ArchitectureSearchConfig]

Custom architecture search configuration

None
**config_overrides Any

Override specific config values (supports nested keys with dots)

{}

Returns:

Type Description
DataFrame

DataFrame containing the learned embeddings

Examples:

Basic usage with defaults

embeddings = learn_embedding_v2(df)

Using a custom config

config = EmbeddingConfig( mode="contrastive", embedding_dim=50, contrastive=ContrastiveConfig(loss_type="triplet", margin=2.0) ) embeddings = learn_embedding_v2(df, config)

embeddings = learn_embedding_v2(df, config, auto_architecture=True)

Quick overrides without config object

embeddings = learn_embedding_v2(df, embedding_dim=20, mode="target", reference_column="category")

Loading from YAML

config = EmbeddingConfig.from_yaml("my_config.yaml") embeddings = learn_embedding_v2(df, config)

learn_embedding_with_model_v2

learn_embedding_with_model_v2(df: DataFrame, config: EmbeddingConfig | None = None, **config_overrides: Any) -> tuple[pd.DataFrame, Row2VecModel]

Modern config-based API for learning embeddings with model artifacts.

This function returns the embeddings along with the trained model, preprocessor, and metadata for serialization purposes.

Parameters:

Name Type Description Default
df DataFrame

Input DataFrame containing the data to embed

required
config EmbeddingConfig | None

Complete embedding configuration. If None, default config is used.

None
**config_overrides Any

Override specific config values

{}

Returns:

Type Description
DataFrame

Tuple of (embeddings, fitted Row2VecModel). The model holds the

Row2VecModel

fitted preprocessor, projector and embedding scaler, and can embed

tuple[DataFrame, Row2VecModel]

new rows with .transform(df).

search_architecture

search_architecture(df: DataFrame, base_config: EmbeddingConfig, search_config: ArchitectureSearchConfig | None = None, target_column: str | None = None) -> tuple[dict[str, Any], ArchitectureSearchResult]

Perform automatic neural architecture search.

This is the main entry point for architecture search functionality.

Parameters:

Name Type Description Default
df DataFrame

Input dataframe for embedding generation

required
base_config EmbeddingConfig

Base embedding configuration

required
search_config ArchitectureSearchConfig | None

Architecture search configuration (uses defaults if None)

None
target_column str | None

Optional target column for supervised evaluation

None

Returns:

Type Description
tuple[dict[str, Any], ArchitectureSearchResult]

Tuple of (best_architecture_dict, full_search_result)

auto_select_dimension

auto_select_dimension(df: DataFrame, config: EmbeddingConfig | None = None, target_column: str | None = None, methods: list[str] | None = None, **selector_kwargs: Any) -> tuple[int, dict[str, Any]]

Convenience function for automatic dimension selection.

Parameters:

Name Type Description Default
df DataFrame

Input dataframe

required
config EmbeddingConfig | None

Base embedding configuration (uses defaults if None)

None
target_column str | None

Optional target column for supervised evaluation

None
methods list[str] | None

List of selection methods to use

None
**selector_kwargs Any

Additional arguments for AutoDimensionSelector

{}

Returns:

Type Description
tuple[int, dict[str, Any]]

Tuple of (optimal_dimension, selection_metadata)

learn_embedding

learn_embedding(df: DataFrame, embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, activation: str = 'relu', early_stopping: bool = True, seed: int = 1305, verbose: bool = False, scale_method: str | None = None, scale_range: tuple[float, float] | None = None, log_level: str = 'INFO', log_file: str | None = None, enable_logging: bool = True, n_neighbors: int = 15, perplexity: float = 30.0, min_dist: float = 0.1, n_iter: int = 1000, similar_pairs: list[tuple[int, int]] | None = None, dissimilar_pairs: list[tuple[int, int]] | None = None, auto_pairs: str | None = None, contrastive_loss: str = 'triplet', margin: float = 1.0, negative_samples: int = 5, aggregate_by_reference: bool = False, config: EmbeddingConfig | None = None) -> pd.DataFrame

Learn an embedding for every row of df.

Parameters:

Name Type Description Default
df DataFrame

Input data with numeric and/or categorical columns.

required
embedding_dim int

Width of the embedding space.

10
mode str

One of "unsupervised" (autoencoder), "target" (supervised encoder), "pca", "tsne", "umap" or "contrastive".

'unsupervised'
reference_column str

Label column. Required for mode="target" and used by auto_pairs="categorical".

None
max_epochs int

Training epoch ceiling for neural modes.

50
batch_size int

Training batch size.

64
dropout_rate float

Dropout applied after each hidden layer.

0.2
hidden_units int | list[int]

Width of the hidden layer, or a list of widths for a multi-layer encoder.

128
activation str

Hidden-layer activation. Previously hardcoded to "relu" no matter what was configured.

'relu'
early_stopping bool

Whether to stop once the monitored loss stalls.

True
seed int

Random seed, applied to Python, NumPy and TensorFlow.

1305
verbose bool

Whether to let the underlying libraries print progress.

False
scale_method str

"none", "minmax", "standard", "l2" or "tanh". The fitted scaler is retained, so a saved model reproduces these values rather than rescaling against whatever batch it is given.

None
scale_range tuple[float, float]

Output range for "minmax".

None
log_level str

Logging level.

'INFO'
log_file str

Path to write logs to.

None
enable_logging bool

Whether to log at all.

True
n_neighbors int

UMAP neighbourhood size.

15
perplexity float

t-SNE perplexity.

30.0
min_dist float

UMAP minimum distance.

0.1
n_iter int

t-SNE iteration count.

1000
similar_pairs list[tuple[int, int]]

Explicit positive pairs, given as positions into df.

None
dissimilar_pairs list[tuple[int, int]]

Explicit negative pairs, given as positions into df.

None
auto_pairs str

"cluster", "neighbors", "categorical" or "random".

None
contrastive_loss str

"triplet" or "contrastive".

'triplet'
margin float

Contrastive loss margin.

1.0
negative_samples int

Negatives drawn per positive pair.

5
aggregate_by_reference bool

mode="target" only. When True, return one row per distinct reference value rather than one row per input row. Defaults to False so target mode, like every other mode, returns a frame indexed by df.index.

False
config EmbeddingConfig

Supplies the preprocessing settings (categorical encoding thresholds, imputation, numeric scaling).

None

Returns:

Type Description
DataFrame

pd.DataFrame: Embeddings in columns embedding_0 through

DataFrame

embedding_{n-1}, indexed by df.index so the result joins

DataFrame

straight back onto df. With aggregate_by_reference=True the

DataFrame

index is the distinct reference values instead.

Raises:

Type Description
ValueError

If the inputs are invalid, or if mode="target" and the reference column contains missing values.

Examples:

>>> import row2vec
>>> df = row2vec.generate_synthetic_data(60)
>>> embeddings = row2vec.learn_embedding(df, mode="pca", embedding_dim=2)
>>> embeddings.shape
(60, 2)

The index is preserved, so the result concatenates cleanly:

>>> list(embeddings.index) == list(df.index)
True

learn_embedding_with_model

learn_embedding_with_model(df: DataFrame, embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, activation: str = 'relu', early_stopping: bool = True, seed: int = 1305, verbose: bool = False, scale_method: str | None = None, scale_range: tuple[float, float] | None = None, log_level: str = 'INFO', log_file: str | None = None, enable_logging: bool = True, n_neighbors: int = 15, perplexity: float = 30.0, min_dist: float = 0.1, n_iter: int = 1000, similar_pairs: list[tuple[int, int]] | None = None, dissimilar_pairs: list[tuple[int, int]] | None = None, auto_pairs: str | None = None, contrastive_loss: str = 'triplet', margin: float = 1.0, negative_samples: int = 5, aggregate_by_reference: bool = False, config: EmbeddingConfig | None = None) -> tuple[pd.DataFrame, Row2VecModel]

Learn embeddings and hand back the fitted model that produced them.

Identical to :func:learn_embedding apart from the return value. A single training pass produces both; the previous implementation ran a second, different training and returned embeddings from one model and the other model itself, at roughly 1.8x the cost.

Parameters:

Name Type Description Default
df DataFrame

Input data.

required
embedding_dim int

Width of the embedding space.

10
mode str

Embedding mode. See :func:learn_embedding.

'unsupervised'
reference_column str

Label column for mode="target".

None
max_epochs int

Training epoch ceiling for neural modes.

50
batch_size int

Training batch size.

64
dropout_rate float

Dropout applied after each hidden layer.

0.2
hidden_units int | list[int]

Hidden layer width, or widths.

128
activation str

Hidden-layer activation.

'relu'
early_stopping bool

Whether to stop once the loss stalls.

True
seed int

Random seed.

1305
verbose bool

Whether to print training progress.

False
scale_method str

Embedding scaling method.

None
scale_range tuple[float, float]

Range for "minmax".

None
log_level str

Logging level.

'INFO'
log_file str

Path to write logs to.

None
enable_logging bool

Whether to log at all.

True
n_neighbors int

UMAP neighbourhood size.

15
perplexity float

t-SNE perplexity.

30.0
min_dist float

UMAP minimum distance.

0.1
n_iter int

t-SNE iteration count.

1000
similar_pairs list[tuple[int, int]]

Explicit positive pairs.

None
dissimilar_pairs list[tuple[int, int]]

Explicit negatives.

None
auto_pairs str

Automatic pairing strategy.

None
contrastive_loss str

"triplet" or "contrastive".

'triplet'
margin float

Contrastive loss margin.

1.0
negative_samples int

Negatives drawn per positive pair.

5
aggregate_by_reference bool

Return one row per reference value.

False
config EmbeddingConfig

Preprocessing settings.

None

Returns:

Type Description
DataFrame

tuple[pd.DataFrame, Row2VecModel]: The training embeddings, and the

Row2VecModel

fitted model. The model embeds new rows with .transform(df) for

tuple[DataFrame, Row2VecModel]

every mode except "tsne", which has no out-of-sample extension.

Examples:

>>> import row2vec
>>> df = row2vec.generate_synthetic_data(60)
>>> embeddings, model = row2vec.learn_embedding_with_model(
...     df, mode="pca", embedding_dim=2
... )
>>> model.transform(df.head(2)).shape
(2, 2)

compare_modes

compare_modes(df: DataFrame, *, target: str | None = None, modes: Sequence[str] | None = None, embedding_dim: int = 2, n_neighbors: int = 10, test_size: float = 0.25, seed: int = 1305, tsne_max_rows: int = 1500, **learn_kwargs: Any) -> pd.DataFrame

Fit several embedding modes on one table and score them side by side.

Parameters:

Name Type Description Default
df DataFrame

The table to embed.

required
target str

A column to predict from the embedding. It is removed from the features for every mode, so no mode sees it as an input, except mode="target", which uses it as its label. Without a target only trustworthiness is scored, and mode="target" is skipped.

None
modes Sequence[str]

Modes to compare. Defaults to every mode that can run ("target" only when target is given).

None
embedding_dim int

Width of every embedding.

2
n_neighbors int

Neighbourhood size for trustworthiness.

10
test_size float

Share of rows held out for scoring.

0.25
seed int

Seed for the split and for every mode.

1305
tsne_max_rows int

t-SNE embeds a seeded random sample of at most this many rows, because its cost grows steeply with the row count (several minutes at a few thousand rows). The sample size is given in its note.

1500
**learn_kwargs Any

Passed to :func:row2vec.learn_embedding_with_model for every mode, e.g. max_epochs=20 to keep neural modes quick. Contrastive mode pairs nearest neighbours (auto_pairs="neighbors") unless you pass auto_pairs or explicit pairs.

{}

Returns:

Type Description
DataFrame

pd.DataFrame: One row per mode, indexed by mode name, with columns

DataFrame

status ("ok", "unavailable", "skipped" or "failed"),

DataFrame

trustworthiness, downstream_score, downstream_metric,

DataFrame

fit_seconds and note. A baseline row gives the same

DataFrame

downstream score on the preprocessed features without any embedding, so

DataFrame

the numbers have something to be compared with. mode="target" is

DataFrame

skipped for a numeric target, which it would treat as one class per

DataFrame

value. Modes that need

DataFrame

TensorFlow are reported as unavailable when it is not installed;

DataFrame

a mode that raises is reported as failed with the error in note.

Raises:

Type Description
ValueError

If target is not a column, has missing values, or a mode is unknown; or if modes asks for "target" without a target.

Examples:

>>> import row2vec
>>> df = row2vec.generate_synthetic_data(120)
>>> report = row2vec.compare_modes(df, target="Country", modes=["pca"])
>>> list(report.index)
['pca', 'baseline']
>>> report.loc["pca", "status"]
'ok'

get_logger

get_logger(name: str = 'row2vec', level: str = 'INFO', log_file: str | Path | None = None, **kwargs: Any) -> Row2VecLogger

Create a Row2Vec logger with standard configuration.

Parameters:

Name Type Description Default
name str

Logger name

'row2vec'
level str

Logging level

'INFO'
log_file str | Path | None

Optional log file path

None
**kwargs Any

Additional arguments for Row2VecLogger

{}

Returns:

Type Description
Row2VecLogger

Configured Row2VecLogger instance

build_adaptive_pipeline

build_adaptive_pipeline(df: DataFrame, target: Series | None = None, config: EmbeddingConfig | None = None, mode: str = 'unsupervised') -> tuple[ColumnTransformer, dict[str, Any]]

Build adaptive preprocessing pipeline for Row2Vec.

This is the main entry point for intelligent pipeline construction. It analyzes the dataset and automatically selects optimal preprocessing strategies based on data characteristics.

Parameters:

Name Type Description Default
df DataFrame

Input dataset

required
target Series

Target variable for supervised preprocessing

None
config EmbeddingConfig

Configuration for preprocessing. If None, intelligent defaults are used.

None
mode str

Embedding mode ("unsupervised", "target", etc.)

'unsupervised'

Returns:

Type Description
tuple[ColumnTransformer, dict[str, Any]]

Tuple[ColumnTransformer, Dict[str, Any]] Preprocessing pipeline and analysis report

Examples:

>>> import row2vec
>>> from row2vec.pipeline_builder import build_adaptive_pipeline
>>> df = row2vec.generate_synthetic_data(60)
>>> pipeline, report = build_adaptive_pipeline(df)
>>> pipeline.fit_transform(df).shape[0]
60
>>> "dataset_shape" in report
True

inspect_model

inspect_model(path: str | Path) -> dict[str, Any]

Read a saved model's manifest without loading the model.

Nothing from the file is deserialised beyond JSON, so this is safe to call on a model you have not decided to trust yet.

Parameters:

Name Type Description Default
path str | Path

A .r2v file.

required

Returns:

Type Description
dict[str, Any]

dict[str, Any]: Format and library versions, mode, embedding dimension,

dict[str, Any]

and the training metadata (columns, dtypes, history, ...).

load_model

load_model(path: str | Path, text_encoder: Callable[[list[str]], Any] | None = None) -> Row2VecModel

Load a model written by :func:save_model.

Loading does not execute code from the file; see the module docstring for exactly what is and is not guaranteed.

Parameters:

Name Type Description Default
path str | Path

The .r2v file (the suffix may be omitted).

required
text_encoder callable

The list[str] -> array hook the model was trained with. Required, and only allowed, for a model trained with a custom text_encoder: the callable is code, so it is not stored in the file.

None

Returns:

Name Type Description
Row2VecModel Row2VecModel

The restored model, ready to transform/predict.

Raises:

Type Description
FileNotFoundError

If the file does not exist.

ModelFormatError

If the file is corrupt, from another format version, names a type that is not on the allow-list, is the old script-and-pickle format, or needs a text_encoder that was not given.

ValueError

If a text_encoder is given for a model that has none.

save_model

save_model(model: Row2VecModel, base_path: str | Path, overwrite: bool = False) -> str

Save a fitted model as a single .r2v file.

Parameters:

Name Type Description Default
model Row2VecModel

The fitted model.

required
base_path str | Path

Where to write it. .r2v is appended unless already present.

required
overwrite bool

Whether to replace an existing file.

False

Returns:

Name Type Description
str str

The path of the file written.

Raises:

Type Description
FileExistsError

If the file exists and overwrite is false.

ValueError

If the model is not fitted.

ModelFormatError

If the model contains something that cannot be saved safely, such as a UMAP model fitted with a custom metric.

Examples:

>>> import tempfile, row2vec
>>> df = row2vec.generate_synthetic_data(60)
>>> _, model = row2vec.learn_embedding_with_model(
...     df, mode="pca", embedding_dim=2, enable_logging=False
... )
>>> with tempfile.TemporaryDirectory() as tmp:
...     path = row2vec.save_model(model, tmp + "/demo")
...     restored = row2vec.load_model(path)
...     path.endswith(".r2v"), restored.predict(df).shape
(True, (60, 2))

train_and_save_model

train_and_save_model(df: DataFrame, base_path: str | Path, embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int = 128, early_stopping: bool = True, seed: int = 1305, verbose: bool = False, scale_method: str | None = None, scale_range: tuple[float, float] | None = None, log_level: str = 'INFO', log_file: str | None = None, enable_logging: bool = True, n_neighbors: int = 15, perplexity: float = 30.0, min_dist: float = 0.1, n_iter: int = 1000, similar_pairs: list[tuple[int, int]] | None = None, dissimilar_pairs: list[tuple[int, int]] | None = None, auto_pairs: str | None = None, negative_samples: int = 5, contrastive_loss: str = 'triplet', margin: float = 1.0, overwrite: bool = False, include_training_history: bool = True) -> tuple[pd.DataFrame, str]

Train a Row2Vec model and save it as a single .r2v file.

This is a convenience function that combines training and saving.

Parameters:

Name Type Description Default
df DataFrame

The input DataFrame containing numeric and categorical features.

required
base_path str | Path

Where to save the model; .r2v is appended unless present.

required
embedding_dim int

The dimensionality of the embedding space.

10
mode str

Embedding method - 'unsupervised' (autoencoder), 'target' (supervised), 'pca' (Principal Component Analysis), 'tsne' (t-SNE), 'umap' (UMAP), or 'contrastive' (contrastive learning).

'unsupervised'
reference_column str

The target column for 'target' mode.

None
max_epochs int

The maximum number of training epochs (neural methods only).

50
batch_size int

The batch size for training (neural methods only).

64
dropout_rate float

The dropout rate for regularization (neural methods only).

0.2
hidden_units Union[int, list[int]]

Hidden layer configuration - single int for one layer or list of ints for multiple layers (neural methods only).

128
early_stopping bool

Whether to use early stopping (neural methods only).

True
seed int

A random seed for reproducibility.

1305
verbose bool

Whether to print training progress.

False
scale_method str

Scaling method for embeddings. Options: 'none', 'minmax', 'standard', 'l2', 'tanh'.

None
scale_range tuple

Range for minmax scaling. Default: (0, 1).

None
log_level str

Logging level ('DEBUG', 'INFO', 'WARNING', 'ERROR').

'INFO'
log_file str

File path for logging output.

None
enable_logging bool

Whether to enable structured logging.

True
n_neighbors int

Number of neighbors for UMAP (default: 15).

15
perplexity float

Perplexity parameter for t-SNE (default: 30.0).

30.0
min_dist float

Minimum distance for UMAP (default: 0.1).

0.1
n_iter int

Number of iterations for t-SNE (default: 1000).

1000
similar_pairs list[tuple[int, int]]

List of (row_idx1, row_idx2) pairs that should have similar embeddings (for contrastive mode).

None
dissimilar_pairs list[tuple[int, int]]

List of (row_idx1, row_idx2) pairs that should have dissimilar embeddings (for contrastive mode).

None
auto_pairs str

Strategy for automatic pair generation. Options: 'cluster' (cluster-based), 'neighbors' (k-NN based), 'categorical' (same category values), 'random' (random sampling).

None
contrastive_loss str

Contrastive loss function. Options: 'triplet', 'contrastive'.

'triplet'
margin float

Margin parameter for contrastive loss functions (default: 1.0).

1.0
negative_samples int

Number of negative samples per positive pair (default: 5).

5
overwrite bool

Whether to overwrite existing model files.

False
include_training_history bool

Whether to include the full training history in the metadata.

True

Returns:

Type Description
tuple[DataFrame, str]

Tuple of (embeddings, path), where path is the .r2v file written.

create_dataframe_schema

create_dataframe_schema(df: DataFrame) -> dict[str, Any]

Create a schema dictionary from a DataFrame for validation purposes.

Parameters:

Name Type Description Default
df DataFrame

DataFrame to analyze

required

Returns:

Type Description
dict[str, Any]

Dictionary containing schema information

generate_synthetic_data

generate_synthetic_data(num_records: int, seed: int = 1305) -> pd.DataFrame

Generates a synthetic DataFrame for demonstration purposes.

Parameters:

Name Type Description Default
num_records int

The number of records to generate.

required
seed int

A random seed for reproducibility.

1305

Returns:

Type Description
DataFrame

pd.DataFrame: A synthetic DataFrame with mixed data types.

validate_dataframe_schema

validate_dataframe_schema(df: DataFrame, expected_schema: dict[str, Any], allow_extra_columns: bool = False, allow_missing_columns: bool = False) -> None

Validate DataFrame schema against expected schema.

Parameters:

Name Type Description Default
df DataFrame

DataFrame to validate

required
expected_schema dict[str, Any]

Expected schema dictionary

required
allow_extra_columns bool

Whether to allow extra columns in df

False
allow_missing_columns bool

Whether to allow missing columns in df

False

Raises:

Type Description
ValueError

If schema validation fails