我正在使用Google colab,并且正在尝试训练卷积神经网络。为了分割大约11,500张图像的数据集,每个数据的形状为63x63x63。我使用了train_test_split
test_split = 0.1
random_state = 42
X_train, X_test, y_train, y_test = train_test_split(triplets, df.label, test_size = test_split, random_state = random_state)
答案 0 :(得分:0)
说明: 由于数据形状为11500x63x63x63,因此数据中大约有3x10 ^ 9个存储位置(实际值为2875540500 500)。通常,一台机器每秒可以执行10 ^ 7〜10 ^ 8条指令。由于python相对较慢,因此我认为google-colab每秒能够执行10 ^ 7条指令,
train_test_split = 3x10 ^ 9/10 ^ 7 = 300秒= 5分钟所需的最短时间
# Building a index array of the input feature
X_index = np.arange(0, 11500)
# Passing index array instead of the big feature matrix
X_train, X_test, y_train, y_test = train_test_split(X_index, df.label, test_size=0.1, random_state=42)
# Extracting the feature matrix using splitted index matrix
X_train = triplets[X_train]
X_test = triplets[X_test]
import timeit
import numpy as np
from sklearn.model_selection import train_test_split
def benchmark(dtypes):
for dtype in dtypes:
print('Benchmark for dtype', dtype, end='\n'+'-'*40+'\n')
X = np.ones((5000, 63, 63, 63), dtype=dtype)
y = np.ones((5000, 1), dtype=dtype)
X_index = np.arange(0, 5000)
start_time = timeit.default_timer()
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.1, random_state=42)
print(f'Time elapsed: {timeit.default_timer()-start_time:.3f}')
start_time = timeit.default_timer()
X_train, X_test, y_train, y_test = train_test_split(X_index, y, test_size=0.1, random_state=42)
X_train = X[X_train]
X_test = X[X_test]
print(f'Time elapsed with indexing: {timeit.default_timer()-start_time:.3f}')
benchmark([np.int8, np.int16, np.int32, np.int64, np.float16, np.float32, np.float64])
Benchmark for dtype <class 'numpy.int8'>
Time elapsed: 0.473
Time elapsed with indexing: 0.304
Benchmark for dtype <class 'numpy.int16'>
Time elapsed: 0.895
Time elapsed with indexing: 0.604
Benchmark for dtype <class 'numpy.int32'>
Time elapsed: 1.792
Time elapsed with indexing: 1.182
Benchmark for dtype <class 'numpy.int64'>
Time elapsed: 2.493
Time elapsed with indexing: 2.398
Benchmark for dtype <class 'numpy.float16'>
Time elapsed: 0.730
Time elapsed with indexing: 0.738
Benchmark for dtype <class 'numpy.float32'>
Time elapsed: 1.904
Time elapsed with indexing: 1.400
Benchmark for dtype <class 'numpy.float64'>
Time elapsed: 5.166
Time elapsed with indexing: 3.076