如何从适合GridSearchCV
中提取最佳管道,以便将其传递给cross_val_predict
?
直接传递fit GridSearchCV
对象会导致cross_val_predict
再次运行整个网格搜索,我只想让最佳管道进行cross_val_predict
评估。
我的自包含代码如下:
from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.svm import SVC
from sklearn.multiclass import OneVsRestClassifier
from sklearn.pipeline import Pipeline
from sklearn.grid_search import GridSearchCV
from sklearn.model_selection import cross_val_predict
from sklearn.model_selection import StratifiedKFold
from sklearn import metrics
# fetch data data
newsgroups = fetch_20newsgroups(remove=('headers', 'footers', 'quotes'), categories=['comp.graphics', 'rec.sport.baseball', 'sci.med'])
X = newsgroups.data
y = newsgroups.target
# setup and run GridSearchCV
wordvect = TfidfVectorizer(analyzer='word', lowercase=True)
classifier = OneVsRestClassifier(SVC(kernel='linear', class_weight='balanced'))
pipeline = Pipeline([('vect', wordvect), ('classifier', classifier)])
scoring = 'f1_weighted'
parameters = {
'vect__min_df': [1, 2],
'vect__max_df': [0.8, 0.9],
'classifier__estimator__C': [0.1, 1, 10]
}
gs_clf = GridSearchCV(pipeline, parameters, n_jobs=8, scoring=scoring, verbose=1)
gs_clf = gs_clf.fit(X, y)
### outputs: Fitting 3 folds for each of 12 candidates, totalling 36 fits
# manually extract the best models from the grid search to re-build the pipeline
best_clf = gs_clf.best_estimator_.named_steps['classifier']
best_vectorizer = gs_clf.best_estimator_.named_steps['vect']
best_pipeline = Pipeline([('best_vectorizer', best_vectorizer), ('classifier', best_clf)])
# passing gs_clf here would run the grind search again inside cross_val_predict
y_predicted = cross_val_predict(pipeline, X, y)
print(metrics.classification_report(y, y_predicted, digits=3))
我目前正在做的是从best_estimator_
手动重建管道。但我的管道通常有更多的步骤,如SVD或PCA,有时我添加或删除步骤并重新运行网格搜索来探索数据。在手动重新构建管道时,必须始终在下面重复该步骤,这很容易出错。
有没有办法直接从适合GridSearchCV
中提取最佳管道,以便我可以将其传递给cross_val_predict
?
答案 0 :(得分:1)
y_predicted = cross_val_predict(gs_clf.best_estimator_, X, y)
工作并返回:
Fitting 3 folds for each of 12 candidates, totalling 36 fits
[Parallel(n_jobs=4)]: Done 36 out of 36 | elapsed: 43.6s finished
precision recall f1-score support
0 0.920 0.911 0.916 584
1 0.894 0.943 0.918 597
2 0.929 0.887 0.908 594
avg / total 0.914 0.914 0.914 1775
[编辑]当我再次尝试传递代码pipeline
(原始管道)时,它返回相同的输出(与传递best_pipeline
一样)。因此,您可以使用Pipeline本身,但我不是100%。
答案 1 :(得分:1)
根据需要为gridsearch对象命名,然后使用fit方法获取结果。您不需要再次交叉验证,因为GridSearchCV基本上是使用不同参数进行交叉验证(FYI,您可以在GridSearchCV中命名自己的cv对象,在sklearn文档中检查GridSearchCV)。
any_name = sklearn.grid_search.GridSearchCV(pipeline, param_grid=parameters)
any_name.fit(X_train, y_train)
以下是我发现的好指南的链接: https://www.civisanalytics.com/blog/workflows-in-python-using-pipeline-and-gridsearchcv-for-more-compact-and-comprehensive-code/
关于SO的第一个答案,希望它有所帮助。