Breaking Changes
- Remove the unused
cache_trainset_representationargument fromArchitectureModule.get_architecture,load_modelandload_model_criterion_config, and thefit_modeargument frominitialize_tabpfn_model. (#1250) - Replace
ClassifierModelSpecs,RegressorModelSpecs, andBaseModelSpecswith the unifiedModelSpecsdataclass. Update imports, constructors, and task-specific type checks; retainnorm_criterionfor legacy regression models without embedded borders. (#1315)
Added
- Add
kv_cache_precision="adaptive". (#1299) - Add a shared
ModelSpecsdataclass for in-memory classification and regression. Regression automatically derives its distribution frommodel.regression_borders; legacy models without embedded borders still requirenorm_criterion. (#1315)
Changed
- Faster modality detection on wide inputs: a numeric array's distinct values are counted for all columns at once instead of per column, a numeric column stored as
objectis recognized withpd.api.types.infer_dtypeinstead of a value-by-value walk, and the nullable-dtype coercion is decided once per dtype. The inferred modalities are unchanged. (#1255) - Building a model from a checkpoint no longer runs torch's random parameter initialisation, which made up most of the per-fit model construction time. (#1257)
SAMPLE_SUBSAMPLING_METHOD="majority_downsample"now corrects the target prior shift it introduces, so predicted probabilities and regression distributions are calibrated to the training data rather than to the downsampled context.predict_proba_batchedandpredict_batchedraise with this sampler; score datasets individually instead. (#1270)- Relabel the README architecture and attention diagrams as TabPFN-3.5. (#1282)
-
examples/finetune_regressor.pynow fine-tunes on the bigger OpenML diamonds dataset. (#1290)
n_preprocessing_jobs > 1now parallelises the per-estimator preprocessing over threads instead of processes, which we have found to be faster. Parallelisation is turned off by default, as previously. (#1295)- Link the TabPFN-3.5 technical report and add its BibTeX entry to the README citation section. (#1296)
-
- Batch estimators. (#1312)
- Estimators fitted with
fit_mode="fit_with_cache"use less host memory and save to smaller files. Predictions are unchanged. (#1323) - Note TabPFN-3.5-Fast alongside TabPFN-3 and TabPFN-3.5 in the README CPU sample-limit description. (#1328)
Fixed
- Fix
FullSupportBarDistributionCDF, quantiles, and sampling to match its half-normal tails while preserving batch shape, device, and dtype. (#1215) - The opt-in built-model cache (
TABPFN_MODEL_CACHE_SIZE) now keys on device and inference precision, so estimators with different settings no longer share one model. (#1219) - numpy
StringDTypearrays are accepted as input at fit and predict, like unicode andobjectarrays (numpy 2.5 or newer). (#1269) - Regression density and likelihood evaluation reports far fewer infinite NLLs for targets in a narrow bucket or in a distribution's tail. (#1288)
- Fixed
FinetunedTabPFNRegressorcomputing the loss against z-scored targets for ensemble members that use a target transform (such as the defaultsafepower); their loss now uses the transformed targets they predict in. (#1290) - MPS support and GQA attention now detect the installed torch version correctly when
torch.__version__is a plain string. (#1301) - Raise a proper validation error when the training inputs have mismatched lengths or are not two-dimensional, so these mistakes are reported as user errors like all other input problems (#1307)
- Fixed TabPFN-3.5 jumping too far when it predicts past the range of a feature, most visibly on a time feature whose values repeat, such as a year column with one row per month. Predictions change for any table with values outside the range seen at fit. (#1310)
TabPFNRegressorno longer fails or returns meaningless predictions when the target's mean is large relative to its spread (e.g. ID-like values around1e10). (#1311)- Allow fitted estimator archives to save supported dataclass-valued initialization parameters such as InferenceConfig. (#1321)
- Forced float16 squashing preprocessing no longer squashes a finite extreme value to zero. (#1322)
- Saving a fitted estimator no longer copies the model weights, so it uses less memory, and saving a
fit_mode="fit_with_cache"estimator no longer briefly removes its caches while another thread may be predicting. (#1325) - Fixed
predictraisingcould not convert string to floatwhen a column was constant or all-missing at fit and held a string at predict. Such a column now gets its fit-time value back at predict, so it carries no information there. (#1329) - Avoid a runtime C compiler requirement in batched v3 and v3.5 regression target encoding on recent PyTorch versions. (#1338)