API Reference¶
This page is generated automatically from the docstrings in the row2vec
package, so it always matches the installed code.
Row2Vec: A library for learning embeddings from tabular data.
This library provides both neural network and classical machine learning approaches for creating vector embeddings from tabular datasets.
ArchitectureSearchConfig
dataclass
¶
ArchitectureSearchConfig(method: str = 'random', max_trials: int = 30, max_time: float | None = 1800, patience: int = 10, min_improvement: float = 0.01, layer_range: tuple[int, int] = (1, 4), max_layers: int = 4, width_options: list[int] = (lambda: [32, 64, 128, 256, 512])(), dropout_options: list[float] = (lambda: [0.0, 0.1, 0.2, 0.3, 0.4, 0.5])(), activation_options: list[str] = (lambda: ['relu', 'elu', 'swish'])(), initial_epochs: int = 10, intermediate_epochs: int = 25, final_epochs: int = 50, top_k_intermediate: int = 10, top_k_final: int = 3, reconstruction_weight: float = 0.4, clustering_weight: float = 0.3, efficiency_weight: float = 0.2, stability_weight: float = 0.1, verbose: bool = True, random_seed: int | None = None, return_full_history: bool = False)
Configuration for automatic neural architecture search.
This class defines the search space, evaluation criteria, and stopping conditions for finding optimal neural network architectures.
ArchitectureSearcher ¶
Main class for performing neural architecture search.
Implements multiple search strategies to find optimal neural network architectures for embedding generation tasks.
search ¶
search(df: DataFrame, base_config: EmbeddingConfig, target_column: str | None = None) -> ArchitectureSearchResult
Perform architecture search on the given dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input dataframe for embedding generation |
required |
base_config
|
EmbeddingConfig
|
Base embedding configuration |
required |
target_column
|
str | None
|
Optional target column for supervised evaluation |
None
|
Returns:
| Type | Description |
|---|---|
ArchitectureSearchResult
|
ArchitectureSearchResult containing the best architecture and metadata |
ArchitectureSearchResult ¶
ArchitectureSearchResult(best_architecture: dict[str, Any], best_score: float, search_history: list[dict[str, Any]], total_time: float, trials_completed: int)
Container for architecture search results.
AutoDimensionSelector ¶
AutoDimensionSelector(methods: list[str] | None = None, performance_weight: float = 0.4, efficiency_weight: float = 0.3, intrinsic_weight: float = 0.3, max_dimension: int | None = None, min_dimension: int = 2, n_trials: int = 5, verbose: bool = True, random_state: int = 1305)
Automatically selects optimal embedding dimensions using multiple strategies.
Combines data-driven analysis, performance optimization, and heuristic rules to determine the best embedding dimension for a given dataset.
Initialize automatic dimension selector.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
methods
|
list[str] | None
|
List of selection methods to use |
None
|
performance_weight
|
float
|
Weight for performance-based selection |
0.4
|
efficiency_weight
|
float
|
Weight for efficiency considerations |
0.3
|
intrinsic_weight
|
float
|
Weight for intrinsic dimensionality estimation |
0.3
|
max_dimension
|
int | None
|
Maximum dimension to consider (auto if None) |
None
|
min_dimension
|
int
|
Minimum dimension to consider |
2
|
n_trials
|
int
|
Number of trials for performance evaluation |
5
|
random_state
|
int
|
Seed for every estimator this selector fits |
1305
|
verbose
|
bool
|
Whether to show selection progress |
True
|
select_dimension ¶
select_dimension(df: DataFrame, config: EmbeddingConfig, target_column: str | None = None, candidate_dims: list[int] | None = None) -> tuple[int, dict[str, Any]]
Select optimal embedding dimension for the given data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input dataframe |
required |
config
|
EmbeddingConfig
|
Base embedding configuration (dimension will be overridden) |
required |
target_column
|
str | None
|
Optional target for supervised evaluation |
None
|
candidate_dims
|
list[int] | None
|
Specific dimensions to evaluate (auto-generated if None) |
None
|
Returns:
| Type | Description |
|---|---|
tuple[int, dict[str, Any]]
|
Tuple of (optimal_dimension, selection_metadata) |
CategoricalAnalyzer ¶
Analyzes categorical data to recommend optimal encoding strategies.
analyze_column ¶
Analyze a categorical column to recommend encoding strategy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
series
|
Series
|
Categorical column to analyze |
required |
target
|
Series
|
Target variable for correlation analysis |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Dict[str, Any] Analysis results and strategy recommendation |
CategoricalEncoder ¶
Bases: BaseEstimator, TransformerMixin
Intelligent categorical encoder with adaptive strategy selection.
This encoder analyzes categorical data characteristics and automatically selects optimal encoding strategies while providing full control for advanced users.
fit ¶
Fit the categorical encoder on training data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
DataFrame
|
Categorical features to encode |
required |
y
|
Series
|
Target variable for supervised encoding strategies |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
CategoricalEncoder
|
Fitted encoder instance |
fit_transform ¶
Fit on X and encode it without leaking the target.
TransformerMixin would give us fit(X, y).transform(X), and
transform deliberately uses the full-data category means - correct
for rows the encoder has never seen, badly leaky for the rows it was
just fitted on. Target-encoded columns therefore take their
cross-fitted values here, matching
:class:sklearn.preprocessing.TargetEncoder.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
DataFrame
|
Categorical features to encode. |
required |
y
|
Series
|
Target for supervised strategies. |
None
|
**fit_params
|
Any
|
Ignored; present for the sklearn signature. |
{}
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: Encoded training features. |
transform ¶
Transform categorical data using fitted encoders.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
DataFrame
|
Categorical data to transform |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame Encoded categorical data |
get_feature_names_out ¶
Get output feature names for transformation.
get_analysis_report ¶
Get detailed analysis report for all columns.
CategoricalEncodingConfig
dataclass
¶
CategoricalEncodingConfig(encoding_strategy: str = 'adaptive', onehot_threshold: int = 20, target_threshold: int = 100, entity_threshold: int = 1000, correlation_threshold: float = 0.1, drop_identifiers: bool = True, min_identifier_rows: int = 50, target_smoothing: float = 1.0, target_noise: float = 0.01, target_cv_folds: int = 5, embedding_dim_ratio: float = 0.5, min_embedding_dim: int = 2, max_embedding_dim: int = 50, entity_epochs: int = 50, entity_batch_size: int = 256, prefer_speed: bool = True, preserve_interpretability: bool = False, enable_feature_selection: bool = False, feature_importance_threshold: float = 0.01, custom_strategies: dict[str, str] = dict(), handle_unknown: str = 'ignore', random_state: int = 42)
Configuration for intelligent categorical encoding strategies.
This class provides comprehensive control over how categorical variables are encoded, with intelligent defaults that automatically select optimal strategies based on data characteristics while allowing expert users to fine-tune every aspect.
encoding_strategy
class-attribute
instance-attribute
¶
Encoding strategy selection. Options: - "adaptive": Automatically selects best strategy based on data analysis - "onehot": One-hot encoding for all categorical features - "target": Target encoding for all categorical features - "entity": Entity embeddings for all categorical features - "ordinal": Ordinal encoding (assumes natural order) - "mixed": Use custom strategies per column (requires custom_strategies)
onehot_threshold
class-attribute
instance-attribute
¶
Use OneHot encoding if cardinality <= this threshold and correlation is low.
target_threshold
class-attribute
instance-attribute
¶
Use target encoding if cardinality is between onehot_threshold and this value.
entity_threshold
class-attribute
instance-attribute
¶
Use entity embeddings if cardinality > target_threshold and <= this value.
correlation_threshold
class-attribute
instance-attribute
¶
Minimum mutual information score to prefer target/entity over onehot.
drop_identifiers
class-attribute
instance-attribute
¶
Drop categorical columns whose every value is distinct (names, ids).
Such a column cannot generalise: no value seen at fit time ever recurs, so
whatever is learned about it describes only the training rows, and every new
row arrives with an unseen value. Encoding it anyway adds noise dimensions.
Applies from min_identifier_rows rows up.
min_identifier_rows
class-attribute
instance-attribute
¶
Fewest rows at which an all-distinct column is judged an identifier.
target_smoothing
class-attribute
instance-attribute
¶
Bayesian smoothing factor for target encoding. Higher values = more smoothing.
target_noise
class-attribute
instance-attribute
¶
Gaussian noise standard deviation added to target encodings to prevent overfitting.
target_cv_folds
class-attribute
instance-attribute
¶
Number of cross-validation folds for target encoding to prevent data leakage.
embedding_dim_ratio
class-attribute
instance-attribute
¶
Embedding dimension as ratio of sqrt(cardinality). Controls embedding size.
min_embedding_dim
class-attribute
instance-attribute
¶
Minimum embedding dimension for entity embeddings.
max_embedding_dim
class-attribute
instance-attribute
¶
Maximum embedding dimension for entity embeddings.
entity_epochs
class-attribute
instance-attribute
¶
Number of training epochs for entity embedding networks.
entity_batch_size
class-attribute
instance-attribute
¶
Batch size for entity embedding training.
prefer_speed
class-attribute
instance-attribute
¶
Whether to prefer faster methods over more accurate but slower ones.
preserve_interpretability
class-attribute
instance-attribute
¶
Whether to prefer interpretable encodings (OneHot/Ordinal) when possible.
enable_feature_selection
class-attribute
instance-attribute
¶
Whether to enable automatic feature selection based on importance.
feature_importance_threshold
class-attribute
instance-attribute
¶
Minimum feature importance score to keep feature (only if enable_feature_selection=True).
custom_strategies
class-attribute
instance-attribute
¶
Custom encoding strategy for specific columns. Format: {column_name: strategy}
handle_unknown
class-attribute
instance-attribute
¶
How to handle unknown categories. Options: 'ignore', 'error', 'infrequent_if_exist'
random_state
class-attribute
instance-attribute
¶
Random state for reproducible results.
EntityEmbeddingTrainer ¶
Trains entity embeddings for high-cardinality categorical features.
fit_column_embedding ¶
fit_column_embedding(series: Series, target: Series | None = None, embedding_dim: int = 10) -> NDArray[Any]
Train entity embeddings for a categorical column.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
series
|
Series
|
Categorical column to embed |
required |
target
|
Series
|
Target variable for supervised embedding |
None
|
embedding_dim
|
int
|
Dimension of embedding vectors |
10
|
Returns:
| Type | Description |
|---|---|
NDArray[Any]
|
np.ndarray Trained embedding matrix of shape (cardinality, embedding_dim) |
TargetEncoder ¶
Implements Bayesian target encoding with cross-validation.
fit_transform ¶
Fit target encoder and transform the series.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
series
|
Series
|
Categorical column to encode |
required |
target
|
Series
|
Target variable |
required |
Returns:
| Type | Description |
|---|---|
Series
|
pd.Series Target-encoded values |
ClassicalConfig
dataclass
¶
ClassicalConfig(n_neighbors: int = 15, min_dist: float = 0.1, perplexity: float = 30.0, n_iter: int = 1000)
Configuration for classical ML dimensionality reduction methods.
ContrastiveConfig
dataclass
¶
ContrastiveConfig(loss_type: str = 'triplet', similar_pairs: list[tuple[int, int]] | None = None, dissimilar_pairs: list[tuple[int, int]] | None = None, auto_pairs: str | None = None, margin: float = 1.0, negative_samples: int = 5)
Configuration for contrastive learning.
EmbeddingConfig
dataclass
¶
EmbeddingConfig(embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, seed: int = 1305, verbose: bool = False, neural: NeuralConfig = NeuralConfig(), classical: ClassicalConfig = ClassicalConfig(), contrastive: ContrastiveConfig = ContrastiveConfig(), scaling: ScalingConfig = ScalingConfig(), logging: LoggingConfig = LoggingConfig(), preprocessing: PreprocessingConfig = PreprocessingConfig())
Complete configuration for embedding learning.
LoggingConfig
dataclass
¶
Configuration for logging and output.
NeuralConfig
dataclass
¶
NeuralConfig(max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, activation: str = 'relu', early_stopping: bool = True)
Configuration for neural network-based embedding methods.
PreprocessingConfig
dataclass
¶
PreprocessingConfig(handle_missing: str = 'auto', numeric_scaling: str = 'standard', categorical_encoding_strategy: str = 'adaptive', categorical_onehot_threshold: int = 20, categorical_target_threshold: int = 100, categorical_entity_threshold: int = 1000, text_columns: list[str] = list(), text_dim: int = 16, text_encoder: Callable[[list[str]], Any] | None = None)
Configuration for data preprocessing including categorical encoding.
ScalingConfig
dataclass
¶
Configuration for embedding scaling/normalization.
AdaptiveImputer ¶
Bases: BaseEstimator
Adaptive imputer that automatically selects and applies appropriate imputation strategies based on data characteristics.
fit ¶
Fit the adaptive imputer to the data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
DataFrame
|
Input DataFrame with potential missing values |
required |
y
|
Any
|
Ignored, present for API compatibility |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
self |
AdaptiveImputer
|
Fitted imputer |
transform ¶
Transform the data by applying imputation strategies.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
DataFrame
|
Input DataFrame with potential missing values |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with missing values imputed |
fit_transform ¶
Fit the imputer and transform the data in one step.
get_imputation_report ¶
Get detailed report about the imputation process.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Dict containing analysis and imputation details |
ImputationConfig
dataclass
¶
ImputationConfig(numeric_strategy: str = 'adaptive', categorical_strategy: str = 'adaptive', prefer_speed: bool = True, missing_threshold: float = 0.7, row_missing_threshold: float = 0.9, knn_neighbors: int = 5, preserve_missing_patterns: bool = False, missing_indicator_suffix: str = '_was_missing', auto_detect_patterns: bool = True, warn_high_missingness: bool = True, categorical_fill_value: str = 'Missing')
Configuration for intelligent missing value imputation strategies.
This class provides comprehensive control over how missing values are handled, with sensible defaults that work well for most datasets while allowing power users to fine-tune every aspect of the imputation process.
numeric_strategy
class-attribute
instance-attribute
¶
Numeric imputation strategy. Options: - "adaptive": Automatically selects best strategy based on missing percentage - "mean": Mean imputation (fastest, good for <10% missing) - "median": Median imputation (robust to outliers, good for 10-30% missing) - "knn": K-nearest neighbors imputation (better for >30% missing) - "iterative": MICE-style iterative imputation (best quality, slowest)
categorical_strategy
class-attribute
instance-attribute
¶
Categorical imputation strategy. Options: - "adaptive": Automatically selects best strategy based on data characteristics - "mode": Most frequent value imputation - "constant": Fill with specified constant value - "missing_category": Create explicit "Missing" category
prefer_speed
class-attribute
instance-attribute
¶
Whether to prefer faster methods over more accurate but slower ones. When True, uses simpler strategies by default. When False, prefers more sophisticated methods even if they take longer.
missing_threshold
class-attribute
instance-attribute
¶
Columns with more than this fraction of missing values will be flagged. Conservative default of 0.7 to avoid dropping useful but sparse columns.
row_missing_threshold
class-attribute
instance-attribute
¶
Rows with more than this fraction of missing values will be flagged. Very conservative default to avoid losing data.
knn_neighbors
class-attribute
instance-attribute
¶
Number of neighbors for KNN imputation. Should be odd to avoid ties.
preserve_missing_patterns
class-attribute
instance-attribute
¶
Whether to preserve missing patterns when they might be informative.
When True, adds binary indicator columns for originally missing values. This is useful when missingness itself carries information (e.g., customers not providing income information might be systematically different).
Example
Original: [1.0, NaN, 3.0] -> After imputation: [1.0, 2.0, 3.0] With preservation: adds column [False, True, False] indicating missingness
missing_indicator_suffix
class-attribute
instance-attribute
¶
Suffix for missing indicator columns when preserve_missing_patterns=True.
auto_detect_patterns
class-attribute
instance-attribute
¶
Whether to automatically analyze missing data patterns and adjust strategies.
warn_high_missingness
class-attribute
instance-attribute
¶
Whether to warn users about columns/rows with high missing percentages.
categorical_fill_value
class-attribute
instance-attribute
¶
Fill value when using 'constant' strategy for categorical data.
MissingPatternAnalyzer ¶
Analyzes missing data patterns to inform imputation strategy selection.
analyze ¶
Analyze missing data patterns in the DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input DataFrame to analyze |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Dict containing analysis results and recommendations |
Row2VecLogger ¶
Row2VecLogger(name: str = 'row2vec', level: str = 'INFO', log_file: str | Path | None = None, include_performance: bool = True, include_memory: bool = True)
Centralized logging system for Row2Vec operations.
Provides structured logging for training progress, debug information, and performance metrics with configurable output formats and levels.
Initialize Row2Vec logger.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Logger name |
'row2vec'
|
level
|
str
|
Logging level (DEBUG, INFO, WARNING, ERROR) |
'INFO'
|
log_file
|
str | Path | None
|
Optional file path for logging output |
None
|
include_performance
|
bool
|
Whether to include performance metrics |
True
|
include_memory
|
bool
|
Whether to include memory usage tracking |
True
|
close ¶
Close and detach the log files this logger opened.
Call this when you are finished with a file-backed logger. Until you do, the file stays open: on Windows that means it cannot be deleted or renamed, and on every platform the handle is held for the life of the process.
Safe to call more than once.
Examples:
start_training ¶
Log training start with configuration details.
log_epoch_metrics ¶
log_epoch_metrics(epoch: int, loss: float, val_loss: float | None = None, additional_metrics: dict[str, float] | None = None) -> None
Log epoch completion with metrics.
end_training ¶
Log training completion with summary.
log_data_preprocessing ¶
Log data preprocessing information.
log_preprocessing_result ¶
log_preprocessing_result(original_shape: tuple[int, int], processed_shape: tuple[int, int], processing_time: float) -> None
Log preprocessing completion.
log_model_architecture ¶
Log model architecture details.
log_performance_warning ¶
Log performance-related warnings.
log_validation_issue ¶
Log validation or data quality issues.
log_debug_info ¶
Log debug information with optional data context.
log_error ¶
Log error with context information.
log_completion ¶
Log completion of embedding generation.
PipelineBuilder ¶
Intelligent pipeline builder that analyzes data and constructs optimal preprocessing pipelines with adaptive strategies.
build_preprocessing_pipeline ¶
build_preprocessing_pipeline(df: DataFrame, target: Series | None = None, mode: str = 'unsupervised') -> tuple[ColumnTransformer, dict[str, Any]]
Build intelligent preprocessing pipeline based on data analysis.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input dataset to analyze |
required |
target
|
Series
|
Target variable for supervised preprocessing |
None
|
mode
|
str
|
Embedding mode that influences preprocessing strategy |
'unsupervised'
|
Returns:
| Type | Description |
|---|---|
tuple[ColumnTransformer, dict[str, Any]]
|
Tuple[ColumnTransformer, Dict[str, Any]] Fitted preprocessing pipeline and analysis report |
get_analysis_report ¶
Get detailed analysis report of the dataset.
get_pipeline_description ¶
Get human-readable description of the constructed pipeline.
ModelFormatError ¶
Bases: ValueError
A saved model is not in a form that can be loaded safely.
Row2VecModel ¶
A fitted row2vec embedder.
Holds every part needed to embed new rows the same way the training rows were embedded: the fitted preprocessing pipeline, the fitted projector, the encoder view for neural modes, and the fitted embedding scaler.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config
|
EmbeddingConfig
|
Configuration to fit with. Defaults to :class: |
None
|
aggregate_by_reference
|
bool
|
|
False
|
Attributes:
| Name | Type | Description |
|---|---|---|
preprocessor_ |
ColumnTransformer
|
The fitted preprocessing pipeline. |
projector_ |
object
|
The fitted PCA/UMAP/t-SNE estimator, or the trained Keras model. |
encoder_ |
object or None
|
For neural modes, the sub-model mapping processed features to the
embedding. |
embedding_scaler_ |
EmbeddingScaler
|
The fitted embedding scaler. |
feature_names_in_ |
list[Hashable]
|
Column labels seen during |
metadata |
object or None
|
Populated by :mod: |
Examples:
>>> import row2vec
>>> df = row2vec.generate_synthetic_data(60)
>>> model = Row2VecModel(row2vec.EmbeddingConfig(embedding_dim=2, mode="pca"))
>>> train = model.fit_transform(df)
>>> list(train.index) == list(df.index)
True
transform reuses the fitted basis rather than refitting:
fit ¶
Fit the preprocessor, projector, and embedding scaler on df.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Training data. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Row2VecModel |
Row2VecModel
|
|
fit_transform ¶
Fit on df and return its embeddings in one training pass.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Training data. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: Embeddings indexed by |
DataFrame
|
|
transform ¶
Embed rows using the state learned during :meth:fit.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Rows to embed. May be a single row. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: Embeddings indexed by |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If the model has not been fitted. |
NotImplementedError
|
If the mode has no out-of-sample extension. |
predict ¶
Embed df with a saved model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Rows to embed. |
required |
validate_schema
|
bool
|
Whether to check |
True
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: Embeddings indexed by |
DataFrame
|
training returned for the same rows. |
validate_input_schema ¶
Check df against the schema recorded when the model was saved.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input to validate. |
required |
strict
|
bool
|
Raise on mismatch rather than returning |
True
|
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
Whether the schema matches. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
to_state ¶
Everything needed to rebuild this model, for pickling.
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
dict[str, Any]: The fitted state, keyed by attribute name. |
from_state
classmethod
¶
Rebuild a fitted model from :meth:to_state output.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state
|
dict[str, Any]
|
A previously saved state dictionary. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Row2VecModel |
Row2VecModel
|
The restored model, ready to |
Row2VecModelMetadata ¶
Row2VecModelMetadata(embedding_dim: int, mode: str, reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, early_stopping: bool = True, seed: int = 1305, scale_method: str | None = None, scale_range: tuple[float, float] | None = None, n_neighbors: int = 15, perplexity: float = 30.0, min_dist: float = 0.1, n_iter: int = 1000, training_history: dict[str, Any] | None = None, final_loss: float | None = None, epochs_trained: int | None = None, training_time: float | None = None, original_columns: list[str] | None = None, preprocessed_feature_names: list[str] | None = None, data_shape: tuple[int, int] | None = None, data_types: dict[str, str] | None = None, expected_schema: dict[str, Any] | None = None)
Row2VecClassifier ¶
Row2VecClassifier(embedding_dim: int = 10, mode: str = 'pca', reference_column: str | None = None, classifier: Any = None, seed: int = 1305, embedding_config: EmbeddingConfig | None = None)
Bases: ClassifierMixin, BaseEstimator
Classify rows using Row2Vec embeddings as features.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
embedding_dim
|
int
|
Width of the embedding space. |
10
|
mode
|
str
|
Embedding mode. See :class: |
'pca'
|
reference_column
|
str
|
Label column for |
None
|
classifier
|
sklearn estimator
|
Downstream classifier.
Cloned before fitting, so the caller's instance is left alone.
Defaults to |
None
|
seed
|
int
|
Random seed. |
1305
|
embedding_config
|
EmbeddingConfig
|
Complete embedding configuration; overrides the individual arguments. |
None
|
Attributes:
| Name | Type | Description |
|---|---|---|
transformer_ |
Row2VecTransformer
|
The fitted embedder. |
classifier_ |
sklearn estimator
|
The fitted downstream classifier. |
classes_ |
ndarray
|
Class labels seen during |
Examples:
>>> import row2vec
>>> from row2vec.sklearn import Row2VecClassifier
>>> df = row2vec.generate_synthetic_data(80)
>>> y = df["Country"]
>>> clf = Row2VecClassifier(embedding_dim=2, mode="pca")
>>> _ = clf.fit(df.drop(columns=["Country"]), y)
>>> len(clf.predict(df.drop(columns=["Country"])))
80
fit ¶
Fit the embedder, then the downstream classifier on its output.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Any
|
Training data. |
required |
y
|
Any
|
Target labels. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Row2VecClassifier |
Row2VecClassifier
|
|
predict ¶
Predict labels for X.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Any
|
Data to classify. |
required |
Returns:
| Type | Description |
|---|---|
ndarray[Any, Any]
|
ndarray of predicted labels. |
predict_proba ¶
Predict class probabilities for X.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Any
|
Data to classify. |
required |
Returns:
| Type | Description |
|---|---|
ndarray[Any, Any]
|
ndarray of shape (n_samples, n_classes). |
Row2VecTransformer ¶
Row2VecTransformer(embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, activation: str = 'relu', seed: int = 1305, config: EmbeddingConfig | None = None)
Bases: TransformerMixin, BaseEstimator
Scikit-learn transformer producing Row2Vec embeddings.
TransformerMixin comes first in the bases deliberately: scikit-learn
reads the transformer tags off the MRO, and with BaseEstimator first
check_estimator refuses to run at all.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
embedding_dim
|
int
|
Width of the embedding space. |
10
|
mode
|
str
|
|
'unsupervised'
|
reference_column
|
str
|
Label column for |
None
|
max_epochs
|
int
|
Training epoch ceiling for neural modes. |
50
|
batch_size
|
int
|
Training batch size. |
64
|
dropout_rate
|
float
|
Dropout after each hidden layer. |
0.2
|
hidden_units
|
int | list[int]
|
Hidden layer width, or widths. |
128
|
activation
|
str
|
Hidden-layer activation. |
'relu'
|
seed
|
int
|
Random seed. |
1305
|
config
|
EmbeddingConfig
|
A complete configuration. When given, the other arguments are ignored. |
None
|
Attributes:
| Name | Type | Description |
|---|---|---|
model_ |
Row2VecModel
|
The fitted model. |
config_ |
EmbeddingConfig
|
The configuration actually used. |
n_features_in_ |
int
|
Number of columns seen during |
feature_names_in_ |
ndarray
|
Column names seen during |
Examples:
>>> import row2vec
>>> from row2vec.sklearn import Row2VecTransformer
>>> df = row2vec.generate_synthetic_data(60)
>>> transformer = Row2VecTransformer(embedding_dim=2, mode="pca")
>>> transformer.fit_transform(df).shape
(60, 2)
A fitted transformer embeds a single unseen row:
fit ¶
Fit the embedding model.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Any
|
Training data. |
required |
y
|
Any
|
Ignored; present for the sklearn signature. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
Row2VecTransformer |
Row2VecTransformer
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
transform ¶
Project data into the fitted embedding space.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Any
|
Data to transform. A single row is fine. |
required |
Returns:
| Type | Description |
|---|---|
ndarray[Any, Any]
|
ndarray of shape (n_samples, embedding_dim). |
fit_transform ¶
Fit and transform in a single training pass.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Any
|
Training data. |
required |
y
|
Any
|
Ignored; present for the sklearn signature. |
None
|
**fit_params
|
Any
|
Ignored; present for the sklearn signature. |
{}
|
Returns:
| Type | Description |
|---|---|
ndarray[Any, Any]
|
ndarray of shape (n_samples, embedding_dim). |
get_feature_names_out ¶
Names of the embedding columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_features
|
ndarray[Any, Any] | None
|
Ignored; present for the sklearn signature. |
None
|
Returns:
| Type | Description |
|---|---|
ndarray[Any, Any]
|
ndarray of str: |
learn_embedding_classical ¶
learn_embedding_classical(df: DataFrame, method: str = 'pca', embedding_dim: int = 10, **overrides: Any) -> pd.DataFrame
Learn classical ML embeddings (PCA, t-SNE, UMAP) with optimized defaults.
learn_embedding_contrastive ¶
Convenience function for contrastive learning with optimized defaults.
learn_embedding_target ¶
learn_embedding_target(df: DataFrame, reference_column: str, embedding_dim: int = 10, **overrides: Any) -> pd.DataFrame
Learn target-based embeddings with optimized defaults.
learn_embedding_unsupervised ¶
learn_embedding_unsupervised(df: DataFrame, embedding_dim: int = 10, **overrides: Any) -> pd.DataFrame
Learn unsupervised embeddings with optimized defaults.
learn_embedding_v2 ¶
learn_embedding_v2(df: DataFrame, config: EmbeddingConfig | None = None, auto_architecture: bool = False, architecture_search_config: Optional[ArchitectureSearchConfig] = None, **config_overrides: Any) -> pd.DataFrame
Modern config-based API for learning embeddings from tabular data.
This is the new recommended API that uses configuration objects instead of long parameter lists. It provides better organization, type safety, and extensibility.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input DataFrame containing the data to embed |
required |
config
|
EmbeddingConfig | None
|
Complete embedding configuration. If None, default config is used. |
None
|
auto_architecture
|
bool
|
Enable automatic neural architecture search for neural modes |
False
|
architecture_search_config
|
Optional[ArchitectureSearchConfig]
|
Custom architecture search configuration |
None
|
**config_overrides
|
Any
|
Override specific config values (supports nested keys with dots) |
{}
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame containing the learned embeddings |
Examples:
Basic usage with defaults¶
embeddings = learn_embedding_v2(df)
Using a custom config¶
config = EmbeddingConfig( mode="contrastive", embedding_dim=50, contrastive=ContrastiveConfig(loss_type="triplet", margin=2.0) ) embeddings = learn_embedding_v2(df, config)
With automatic architecture search¶
embeddings = learn_embedding_v2(df, config, auto_architecture=True)
Quick overrides without config object¶
embeddings = learn_embedding_v2(df, embedding_dim=20, mode="target", reference_column="category")
Loading from YAML¶
config = EmbeddingConfig.from_yaml("my_config.yaml") embeddings = learn_embedding_v2(df, config)
learn_embedding_with_model_v2 ¶
learn_embedding_with_model_v2(df: DataFrame, config: EmbeddingConfig | None = None, **config_overrides: Any) -> tuple[pd.DataFrame, Row2VecModel]
Modern config-based API for learning embeddings with model artifacts.
This function returns the embeddings along with the trained model, preprocessor, and metadata for serialization purposes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input DataFrame containing the data to embed |
required |
config
|
EmbeddingConfig | None
|
Complete embedding configuration. If None, default config is used. |
None
|
**config_overrides
|
Any
|
Override specific config values |
{}
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Tuple of (embeddings, fitted Row2VecModel). The model holds the |
Row2VecModel
|
fitted preprocessor, projector and embedding scaler, and can embed |
tuple[DataFrame, Row2VecModel]
|
new rows with |
search_architecture ¶
search_architecture(df: DataFrame, base_config: EmbeddingConfig, search_config: ArchitectureSearchConfig | None = None, target_column: str | None = None) -> tuple[dict[str, Any], ArchitectureSearchResult]
Perform automatic neural architecture search.
This is the main entry point for architecture search functionality.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input dataframe for embedding generation |
required |
base_config
|
EmbeddingConfig
|
Base embedding configuration |
required |
search_config
|
ArchitectureSearchConfig | None
|
Architecture search configuration (uses defaults if None) |
None
|
target_column
|
str | None
|
Optional target column for supervised evaluation |
None
|
Returns:
| Type | Description |
|---|---|
tuple[dict[str, Any], ArchitectureSearchResult]
|
Tuple of (best_architecture_dict, full_search_result) |
auto_select_dimension ¶
auto_select_dimension(df: DataFrame, config: EmbeddingConfig | None = None, target_column: str | None = None, methods: list[str] | None = None, **selector_kwargs: Any) -> tuple[int, dict[str, Any]]
Convenience function for automatic dimension selection.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input dataframe |
required |
config
|
EmbeddingConfig | None
|
Base embedding configuration (uses defaults if None) |
None
|
target_column
|
str | None
|
Optional target column for supervised evaluation |
None
|
methods
|
list[str] | None
|
List of selection methods to use |
None
|
**selector_kwargs
|
Any
|
Additional arguments for AutoDimensionSelector |
{}
|
Returns:
| Type | Description |
|---|---|
tuple[int, dict[str, Any]]
|
Tuple of (optimal_dimension, selection_metadata) |
learn_embedding ¶
learn_embedding(df: DataFrame, embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, activation: str = 'relu', early_stopping: bool = True, seed: int = 1305, verbose: bool = False, scale_method: str | None = None, scale_range: tuple[float, float] | None = None, log_level: str = 'INFO', log_file: str | None = None, enable_logging: bool = True, n_neighbors: int = 15, perplexity: float = 30.0, min_dist: float = 0.1, n_iter: int = 1000, similar_pairs: list[tuple[int, int]] | None = None, dissimilar_pairs: list[tuple[int, int]] | None = None, auto_pairs: str | None = None, contrastive_loss: str = 'triplet', margin: float = 1.0, negative_samples: int = 5, aggregate_by_reference: bool = False, config: EmbeddingConfig | None = None) -> pd.DataFrame
Learn an embedding for every row of df.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input data with numeric and/or categorical columns. |
required |
embedding_dim
|
int
|
Width of the embedding space. |
10
|
mode
|
str
|
One of |
'unsupervised'
|
reference_column
|
str
|
Label column. Required for
|
None
|
max_epochs
|
int
|
Training epoch ceiling for neural modes. |
50
|
batch_size
|
int
|
Training batch size. |
64
|
dropout_rate
|
float
|
Dropout applied after each hidden layer. |
0.2
|
hidden_units
|
int | list[int]
|
Width of the hidden layer, or a list of widths for a multi-layer encoder. |
128
|
activation
|
str
|
Hidden-layer activation. Previously hardcoded to
|
'relu'
|
early_stopping
|
bool
|
Whether to stop once the monitored loss stalls. |
True
|
seed
|
int
|
Random seed, applied to Python, NumPy and TensorFlow. |
1305
|
verbose
|
bool
|
Whether to let the underlying libraries print progress. |
False
|
scale_method
|
str
|
|
None
|
scale_range
|
tuple[float, float]
|
Output range for
|
None
|
log_level
|
str
|
Logging level. |
'INFO'
|
log_file
|
str
|
Path to write logs to. |
None
|
enable_logging
|
bool
|
Whether to log at all. |
True
|
n_neighbors
|
int
|
UMAP neighbourhood size. |
15
|
perplexity
|
float
|
t-SNE perplexity. |
30.0
|
min_dist
|
float
|
UMAP minimum distance. |
0.1
|
n_iter
|
int
|
t-SNE iteration count. |
1000
|
similar_pairs
|
list[tuple[int, int]]
|
Explicit positive
pairs, given as positions into |
None
|
dissimilar_pairs
|
list[tuple[int, int]]
|
Explicit negative
pairs, given as positions into |
None
|
auto_pairs
|
str
|
|
None
|
contrastive_loss
|
str
|
|
'triplet'
|
margin
|
float
|
Contrastive loss margin. |
1.0
|
negative_samples
|
int
|
Negatives drawn per positive pair. |
5
|
aggregate_by_reference
|
bool
|
|
False
|
config
|
EmbeddingConfig
|
Supplies the preprocessing settings (categorical encoding thresholds, imputation, numeric scaling). |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: Embeddings in columns |
DataFrame
|
|
DataFrame
|
straight back onto |
DataFrame
|
index is the distinct reference values instead. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the inputs are invalid, or if |
Examples:
>>> import row2vec
>>> df = row2vec.generate_synthetic_data(60)
>>> embeddings = row2vec.learn_embedding(df, mode="pca", embedding_dim=2)
>>> embeddings.shape
(60, 2)
The index is preserved, so the result concatenates cleanly:
learn_embedding_with_model ¶
learn_embedding_with_model(df: DataFrame, embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int | list[int] = 128, activation: str = 'relu', early_stopping: bool = True, seed: int = 1305, verbose: bool = False, scale_method: str | None = None, scale_range: tuple[float, float] | None = None, log_level: str = 'INFO', log_file: str | None = None, enable_logging: bool = True, n_neighbors: int = 15, perplexity: float = 30.0, min_dist: float = 0.1, n_iter: int = 1000, similar_pairs: list[tuple[int, int]] | None = None, dissimilar_pairs: list[tuple[int, int]] | None = None, auto_pairs: str | None = None, contrastive_loss: str = 'triplet', margin: float = 1.0, negative_samples: int = 5, aggregate_by_reference: bool = False, config: EmbeddingConfig | None = None) -> tuple[pd.DataFrame, Row2VecModel]
Learn embeddings and hand back the fitted model that produced them.
Identical to :func:learn_embedding apart from the return value. A single
training pass produces both; the previous implementation ran a second,
different training and returned embeddings from one model and the other
model itself, at roughly 1.8x the cost.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input data. |
required |
embedding_dim
|
int
|
Width of the embedding space. |
10
|
mode
|
str
|
Embedding mode. See :func: |
'unsupervised'
|
reference_column
|
str
|
Label column for |
None
|
max_epochs
|
int
|
Training epoch ceiling for neural modes. |
50
|
batch_size
|
int
|
Training batch size. |
64
|
dropout_rate
|
float
|
Dropout applied after each hidden layer. |
0.2
|
hidden_units
|
int | list[int]
|
Hidden layer width, or widths. |
128
|
activation
|
str
|
Hidden-layer activation. |
'relu'
|
early_stopping
|
bool
|
Whether to stop once the loss stalls. |
True
|
seed
|
int
|
Random seed. |
1305
|
verbose
|
bool
|
Whether to print training progress. |
False
|
scale_method
|
str
|
Embedding scaling method. |
None
|
scale_range
|
tuple[float, float]
|
Range for |
None
|
log_level
|
str
|
Logging level. |
'INFO'
|
log_file
|
str
|
Path to write logs to. |
None
|
enable_logging
|
bool
|
Whether to log at all. |
True
|
n_neighbors
|
int
|
UMAP neighbourhood size. |
15
|
perplexity
|
float
|
t-SNE perplexity. |
30.0
|
min_dist
|
float
|
UMAP minimum distance. |
0.1
|
n_iter
|
int
|
t-SNE iteration count. |
1000
|
similar_pairs
|
list[tuple[int, int]]
|
Explicit positive pairs. |
None
|
dissimilar_pairs
|
list[tuple[int, int]]
|
Explicit negatives. |
None
|
auto_pairs
|
str
|
Automatic pairing strategy. |
None
|
contrastive_loss
|
str
|
|
'triplet'
|
margin
|
float
|
Contrastive loss margin. |
1.0
|
negative_samples
|
int
|
Negatives drawn per positive pair. |
5
|
aggregate_by_reference
|
bool
|
Return one row per reference value. |
False
|
config
|
EmbeddingConfig
|
Preprocessing settings. |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
tuple[pd.DataFrame, Row2VecModel]: The training embeddings, and the |
Row2VecModel
|
fitted model. The model embeds new rows with |
tuple[DataFrame, Row2VecModel]
|
every mode except |
Examples:
compare_modes ¶
compare_modes(df: DataFrame, *, target: str | None = None, modes: Sequence[str] | None = None, embedding_dim: int = 2, n_neighbors: int = 10, test_size: float = 0.25, seed: int = 1305, tsne_max_rows: int = 1500, **learn_kwargs: Any) -> pd.DataFrame
Fit several embedding modes on one table and score them side by side.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
The table to embed. |
required |
target
|
str
|
A column to predict from the embedding. It is
removed from the features for every mode, so no mode sees it as an
input, except |
None
|
modes
|
Sequence[str]
|
Modes to compare. Defaults to every
mode that can run ( |
None
|
embedding_dim
|
int
|
Width of every embedding. |
2
|
n_neighbors
|
int
|
Neighbourhood size for trustworthiness. |
10
|
test_size
|
float
|
Share of rows held out for scoring. |
0.25
|
seed
|
int
|
Seed for the split and for every mode. |
1305
|
tsne_max_rows
|
int
|
t-SNE embeds a seeded random sample of at most this
many rows, because its cost grows steeply with the row count (several
minutes at a few thousand rows). The sample size is given in its
|
1500
|
**learn_kwargs
|
Any
|
Passed to :func: |
{}
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: One row per mode, indexed by mode name, with columns |
DataFrame
|
|
DataFrame
|
|
DataFrame
|
|
DataFrame
|
downstream score on the preprocessed features without any embedding, so |
DataFrame
|
the numbers have something to be compared with. |
DataFrame
|
|
DataFrame
|
value. Modes that need |
DataFrame
|
TensorFlow are reported as |
DataFrame
|
a mode that raises is reported as |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Examples:
get_logger ¶
get_logger(name: str = 'row2vec', level: str = 'INFO', log_file: str | Path | None = None, **kwargs: Any) -> Row2VecLogger
Create a Row2Vec logger with standard configuration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
Logger name |
'row2vec'
|
level
|
str
|
Logging level |
'INFO'
|
log_file
|
str | Path | None
|
Optional log file path |
None
|
**kwargs
|
Any
|
Additional arguments for Row2VecLogger |
{}
|
Returns:
| Type | Description |
|---|---|
Row2VecLogger
|
Configured Row2VecLogger instance |
build_adaptive_pipeline ¶
build_adaptive_pipeline(df: DataFrame, target: Series | None = None, config: EmbeddingConfig | None = None, mode: str = 'unsupervised') -> tuple[ColumnTransformer, dict[str, Any]]
Build adaptive preprocessing pipeline for Row2Vec.
This is the main entry point for intelligent pipeline construction. It analyzes the dataset and automatically selects optimal preprocessing strategies based on data characteristics.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input dataset |
required |
target
|
Series
|
Target variable for supervised preprocessing |
None
|
config
|
EmbeddingConfig
|
Configuration for preprocessing. If None, intelligent defaults are used. |
None
|
mode
|
str
|
Embedding mode ("unsupervised", "target", etc.) |
'unsupervised'
|
Returns:
| Type | Description |
|---|---|
tuple[ColumnTransformer, dict[str, Any]]
|
Tuple[ColumnTransformer, Dict[str, Any]] Preprocessing pipeline and analysis report |
Examples:
inspect_model ¶
Read a saved model's manifest without loading the model.
Nothing from the file is deserialised beyond JSON, so this is safe to call on a model you have not decided to trust yet.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str | Path
|
A |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
dict[str, Any]: Format and library versions, mode, embedding dimension, |
dict[str, Any]
|
and the training metadata (columns, dtypes, history, ...). |
load_model ¶
load_model(path: str | Path, text_encoder: Callable[[list[str]], Any] | None = None) -> Row2VecModel
Load a model written by :func:save_model.
Loading does not execute code from the file; see the module docstring for exactly what is and is not guaranteed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str | Path
|
The |
required |
text_encoder
|
callable
|
The |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
Row2VecModel |
Row2VecModel
|
The restored model, ready to |
Raises:
| Type | Description |
|---|---|
FileNotFoundError
|
If the file does not exist. |
ModelFormatError
|
If the file is corrupt, from another format version,
names a type that is not on the allow-list, is the old
script-and-pickle format, or needs a |
ValueError
|
If a |
save_model ¶
Save a fitted model as a single .r2v file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
Row2VecModel
|
The fitted model. |
required |
base_path
|
str | Path
|
Where to write it. |
required |
overwrite
|
bool
|
Whether to replace an existing file. |
False
|
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
The path of the file written. |
Raises:
| Type | Description |
|---|---|
FileExistsError
|
If the file exists and |
ValueError
|
If the model is not fitted. |
ModelFormatError
|
If the model contains something that cannot be saved safely, such as a UMAP model fitted with a custom metric. |
Examples:
>>> import tempfile, row2vec
>>> df = row2vec.generate_synthetic_data(60)
>>> _, model = row2vec.learn_embedding_with_model(
... df, mode="pca", embedding_dim=2, enable_logging=False
... )
>>> with tempfile.TemporaryDirectory() as tmp:
... path = row2vec.save_model(model, tmp + "/demo")
... restored = row2vec.load_model(path)
... path.endswith(".r2v"), restored.predict(df).shape
(True, (60, 2))
train_and_save_model ¶
train_and_save_model(df: DataFrame, base_path: str | Path, embedding_dim: int = 10, mode: str = 'unsupervised', reference_column: str | None = None, max_epochs: int = 50, batch_size: int = 64, dropout_rate: float = 0.2, hidden_units: int = 128, early_stopping: bool = True, seed: int = 1305, verbose: bool = False, scale_method: str | None = None, scale_range: tuple[float, float] | None = None, log_level: str = 'INFO', log_file: str | None = None, enable_logging: bool = True, n_neighbors: int = 15, perplexity: float = 30.0, min_dist: float = 0.1, n_iter: int = 1000, similar_pairs: list[tuple[int, int]] | None = None, dissimilar_pairs: list[tuple[int, int]] | None = None, auto_pairs: str | None = None, negative_samples: int = 5, contrastive_loss: str = 'triplet', margin: float = 1.0, overwrite: bool = False, include_training_history: bool = True) -> tuple[pd.DataFrame, str]
Train a Row2Vec model and save it as a single .r2v file.
This is a convenience function that combines training and saving.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
The input DataFrame containing numeric and categorical features. |
required |
base_path
|
str | Path
|
Where to save the model; |
required |
embedding_dim
|
int
|
The dimensionality of the embedding space. |
10
|
mode
|
str
|
Embedding method - 'unsupervised' (autoencoder), 'target' (supervised), 'pca' (Principal Component Analysis), 'tsne' (t-SNE), 'umap' (UMAP), or 'contrastive' (contrastive learning). |
'unsupervised'
|
reference_column
|
str
|
The target column for 'target' mode. |
None
|
max_epochs
|
int
|
The maximum number of training epochs (neural methods only). |
50
|
batch_size
|
int
|
The batch size for training (neural methods only). |
64
|
dropout_rate
|
float
|
The dropout rate for regularization (neural methods only). |
0.2
|
hidden_units
|
Union[int, list[int]]
|
Hidden layer configuration - single int for one layer or list of ints for multiple layers (neural methods only). |
128
|
early_stopping
|
bool
|
Whether to use early stopping (neural methods only). |
True
|
seed
|
int
|
A random seed for reproducibility. |
1305
|
verbose
|
bool
|
Whether to print training progress. |
False
|
scale_method
|
str
|
Scaling method for embeddings. Options: 'none', 'minmax', 'standard', 'l2', 'tanh'. |
None
|
scale_range
|
tuple
|
Range for minmax scaling. Default: (0, 1). |
None
|
log_level
|
str
|
Logging level ('DEBUG', 'INFO', 'WARNING', 'ERROR'). |
'INFO'
|
log_file
|
str
|
File path for logging output. |
None
|
enable_logging
|
bool
|
Whether to enable structured logging. |
True
|
n_neighbors
|
int
|
Number of neighbors for UMAP (default: 15). |
15
|
perplexity
|
float
|
Perplexity parameter for t-SNE (default: 30.0). |
30.0
|
min_dist
|
float
|
Minimum distance for UMAP (default: 0.1). |
0.1
|
n_iter
|
int
|
Number of iterations for t-SNE (default: 1000). |
1000
|
similar_pairs
|
list[tuple[int, int]]
|
List of (row_idx1, row_idx2) pairs that should have similar embeddings (for contrastive mode). |
None
|
dissimilar_pairs
|
list[tuple[int, int]]
|
List of (row_idx1, row_idx2) pairs that should have dissimilar embeddings (for contrastive mode). |
None
|
auto_pairs
|
str
|
Strategy for automatic pair generation. Options: 'cluster' (cluster-based), 'neighbors' (k-NN based), 'categorical' (same category values), 'random' (random sampling). |
None
|
contrastive_loss
|
str
|
Contrastive loss function. Options: 'triplet', 'contrastive'. |
'triplet'
|
margin
|
float
|
Margin parameter for contrastive loss functions (default: 1.0). |
1.0
|
negative_samples
|
int
|
Number of negative samples per positive pair (default: 5). |
5
|
overwrite
|
bool
|
Whether to overwrite existing model files. |
False
|
include_training_history
|
bool
|
Whether to include the full training history in the metadata. |
True
|
Returns:
| Type | Description |
|---|---|
tuple[DataFrame, str]
|
Tuple of (embeddings, path), where |
create_dataframe_schema ¶
Create a schema dictionary from a DataFrame for validation purposes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
DataFrame to analyze |
required |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Dictionary containing schema information |
generate_synthetic_data ¶
Generates a synthetic DataFrame for demonstration purposes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
num_records
|
int
|
The number of records to generate. |
required |
seed
|
int
|
A random seed for reproducibility. |
1305
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
pd.DataFrame: A synthetic DataFrame with mixed data types. |
validate_dataframe_schema ¶
validate_dataframe_schema(df: DataFrame, expected_schema: dict[str, Any], allow_extra_columns: bool = False, allow_missing_columns: bool = False) -> None
Validate DataFrame schema against expected schema.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
DataFrame to validate |
required |
expected_schema
|
dict[str, Any]
|
Expected schema dictionary |
required |
allow_extra_columns
|
bool
|
Whether to allow extra columns in df |
False
|
allow_missing_columns
|
bool
|
Whether to allow missing columns in df |
False
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If schema validation fails |