Add manual task iteration tutorial - #788
Conversation
Codecov Report
@@ Coverage Diff @@
## develop #788 +/- ##
==========================================
+ Coverage 87.69% 88.2% +0.51%
==========================================
Files 36 36
Lines 4200 4452 +252
==========================================
+ Hits 3683 3927 +244
- Misses 517 525 +8
Continue to review full report at Codecov.
|
| import openml | ||
|
|
||
| #################################################################################################### | ||
| task_id = 233 |
There was a problem hiding this comment.
maybe add small comment which task this is and why (few observations, well-known iris, ... ?)
There was a problem hiding this comment.
Great idea, thanks I lot, I just added that explanation.
| # Task ``233`` is a simple task using the holdout estimation procedure and therefore has only a | ||
| # single repeat, a single fold and a single sample size: | ||
|
|
||
| print(n_repeats, n_folds, n_samples) |
There was a problem hiding this comment.
use logging instead of print (?) also, little bit more verbose would be good:
logging.info('dataset %s, %d repeats, %d folds, %d samples' % (task.get_dataset.name, n_repeats, n_folds, n_samples))
There was a problem hiding this comment.
I made this more verbose.
|
|
||
| #################################################################################################### | ||
| # We can now retrieve the train/test split for this combination of repeats, folds and number of | ||
| # samples (indexing is zero-based): |
There was a problem hiding this comment.
maybe add that usually we do a loop around this, but only neglect this since we have a single repeat / fold
| sample=0, | ||
| ) | ||
|
|
||
| print(train_indices.shape, train_indices.dtype) |
| X_test = X.loc[test_indices] | ||
| y_test = y[test_indices] | ||
|
|
||
| print(X_train.shape, y_train.shape, X_test.shape, y_test.shape) |
| print(X_train.shape, y_train.shape, X_test.shape, y_test.shape) | ||
|
|
||
| #################################################################################################### | ||
| # Obviously, we can also retrieve cross-validation versions of the dataset used in task ``233``: |
There was a problem hiding this comment.
mention that we require 1 for loop over folds
There was a problem hiding this comment.
I added a loop here and for the other two as well.
| print(n_repeats, n_folds, n_samples) | ||
|
|
||
| #################################################################################################### | ||
| # And also versions with multiple repeats: |
| print(n_repeats, n_folds, n_samples) | ||
|
|
||
| #################################################################################################### | ||
| # And finally a task based on learning curves: |
| api_call += "/flow/%s" % ','.join([str(int(i)) for i in flow]) | ||
| if uploader is not None: | ||
| api_call += "/uploader/%s" % ','.join([str(int(i)) for i in uploader]) | ||
| if run is not None: |
There was a problem hiding this comment.
can you also add a unit test that filters based on runs?
There was a problem hiding this comment.
Sorry, this should not have been part of this PR, I will do a new PR.
janvanrijn
left a comment
There was a problem hiding this comment.
looks good. I could only find some very small things.
No description provided.