Skip to main content

Evaluate

talos.Evaluate scores a trained model from a completed Scan or RunResult on caller-supplied held-out observations. Import it from talos. Evaluate(scan_object) stores the run; its .evaluate() method selects one fitted model and scores subsets without retraining.

Prerequisites are recoverable trained models, their installed framework backend, and held-out features and targets with matching sample counts. Classification returns F1; continuous tasks return MAE. The metric argument selects the model from scan results; it does not change the held-out scoring formula.

Examples below use the held-out Iris setup in Scan → Minimal Example. Run that setup first; it defines scan_object, p, input_model, x, y, x_val, y_val, x_test, and y_test.

from talos import Evaluate

# create the evaluate object
e = Evaluate(scan_object)

# perform the evaluation
scores = e.evaluate(x_test, y_test, task='multi_class', metric='val_loss',
asc=True, average='macro', folds=3, seed=17)

NOTE: It's very important to save part of your data for evaluation, and keep it completely separated from the data you use for the actual experiment. Choose an evaluation fraction suited to the dataset; the shared example reserves 20%. These folds score the same fitted model on held-out subsets; they do not retrain it.

Interface​

The method signature is evaluate(x, y, task, metric, model_id=None, folds=5, shuffle=True, asc=False, saved=False, custom_objects=None, multi_input=False, print_out=False, average=None, model_factory=None, seed=None). It returns a list of folds floating-point scores. Evaluate holds .scan_object and its .data table but does not add score columns itself.

Arguments​

ParameterDefaultDescription
x, yrequiredHeld-out features and truth labels, aligned by row.
taskrequiredClassification or continuous task, as described below.
metricrequiredScan result column used to select the fitted model.
model_idNoneExplicit result-row index; otherwise choose the best metric value.
folds5Number of held-out scoring subsets.
shuffleTrueShuffle rows before splitting.
ascFalseUse True to minimize the selection metric.
savedFalseLoad the selected model from persisted artifacts.
custom_objectsNoneKeras objects required for model reconstruction.
multi_inputFalseSet True for a list of feature arrays.
print_outFalsePrint the mean and standard deviation of fold scores.
averageNoneTask default F1 averaging; may be binary, micro, macro, samples, or weighted where supported by scikit-learn.
model_factoryNoneTorch reconstruction factory when needed.
seedNoneSeed for shuffled fold assignment.

The above arguments are for the evaluate attribute of the Evaluate object.

Score semantics​

TaskTarget/output handlingScore and units
binaryScalar output thresholded at 0.5, or two-output predictions reduced with argmax.F1 with binary averaging by default, from 0 to 1.
multi_class / multiclassOne-hot truth and predicted class distributions reduced with argmax.Macro F1 by default, from 0 to 1.
multi_labelOne-hot truth uses multiclass scoring; independent multi-hot truth uses per-label 0.5 thresholds.Macro F1 by default, from 0 to 1.
multilabel / multi_label_independentIndependent per-label 0.5 thresholds.Macro F1 by default, from 0 to 1.
continuous / regressionRaw predictions.Mean absolute error, in the target's units.

List or dictionary targets are scored output by output, then averaged for each subset. Every row enters exactly one subset, including a final uneven subset. Fold means are an unweighted mean of subset scores; for uneven folds this can differ from one score over all held-out rows.

folds must be an integer from 1 through the number of held-out rows. Invalid folds and unknown tasks raise ValueError; list feature inputs without multi_input=True raise TypeError. Missing metrics, unavailable saved models and shape errors propagate from model selection, restoration or scikit-learn. No leakage check can establish that the caller's held-out data was never used for training.

Evaluate several candidates​

scan_object.evaluate_models(x_val, y_val, task, n_models=10, metric='val_acc', folds=5, shuffle=True, asc=False, saved=False, custom_objects=None, average=None, model_factory=None, multi_input=None, seed=None) evaluates the top candidates selected by metric. It returns None and adds eval_f1score_mean/eval_f1score_std, or eval_mae_mean/eval_mae_std, to .data. Unselected rows receive missing values. With multi_input=None, a list of feature arrays is detected automatically.

These added columns are an in-memory table update; they do not rewrite the completed run's result files. Persist an exported table yourself when those post-run scores must be retained.

Use Predict for inference with an explicit selection metric, AutoPredict for candidate evaluation and winner prediction, or Deploy for packaging.