将节点添加到Hadoop集群时,为什么吞吐量和平均io速率变慢?

时间:2019-06-20 17:17:10

标签: hadoop io mapreduce benchmarking throughput

因此,我在集群上运行了TestDFSIO来查看吞吐量和平均读写操作速率。 我做4测试: 4个文件,每个256 MB(总计1 GB) 2个文件,每个256 MB(总计512 MB) 2个文件,每个128 MB(总计256 MB) 1个文件50 MB(总共50 MB)

我在单节点到5节点的hadoop集群上运行它们。具有256 MB的块大小,并且每个节点具有不同的复制(单节点= 1复制,2节点= 2复制,依此类推)。

这是1 GB数据测试的测试结果 1个节点

----- TestDFSIO ----- : write
           Date & time: Thu Jun 20 11:38:21 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 8.503288381053611
Average IO rate mb/sec: 8.507380485534668
 IO rate std deviation: 0.18595730311606032
    Test exec time sec: 84.876

----- TestDFSIO ----- : read
           Date & time: Thu Jun 20 11:39:52 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 14.351786965662228
Average IO rate mb/sec: 14.422638893127441
 IO rate std deviation: 1.0515649052955383
    Test exec time sec: 61.371

2 node
----- TestDFSIO ----- : write
           Date & time: Thu Jun 20 11:15:52 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 2.557167936510315
Average IO rate mb/sec: 2.5574562549591064
 IO rate std deviation: 0.027311795003682558
    Test exec time sec: 150.506

----- TestDFSIO ----- : read
           Date & time: Thu Jun 20 11:18:04 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 9.567321617101587
Average IO rate mb/sec: 9.673456192016602
 IO rate std deviation: 1.0593562755825534
    Test exec time sec: 79.333

3 node
----- TestDFSIO ----- : write
           Date & time: Thu Jun 20 10:42:47 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 2.343067129788529
Average IO rate mb/sec: 2.3866918087005615
 IO rate std deviation: 0.3233444726530288
    Test exec time sec: 167.593

----- TestDFSIO ----- : read
           Date & time: Thu Jun 20 10:47:33 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 11.901164547546546
Average IO rate mb/sec: 12.255699157714844
 IO rate std deviation: 2.2415787547598667
    Test exec time sec: 69.29

4 node 
----- TestDFSIO ----- : write
           Date & time: Thu Jun 20 10:23:19 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 1.6539390885245053
Average IO rate mb/sec: 1.6625666618347168
 IO rate std deviation: 0.12093049037575003
    Test exec time sec: 205.164

----- TestDFSIO ----- : read
           Date & time: Thu Jun 20 10:25:23 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 19.842653954966476
Average IO rate mb/sec: 20.02923583984375
 IO rate std deviation: 1.9719328195872965
    Test exec time sec: 57.25

5 node
----- TestDFSIO ----- : write
           Date & time: Thu Jun 13 12:50:12 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 1.5617159964556366
Average IO rate mb/sec: 1.573684573173523
 IO rate std deviation: 0.14426118715726127
    Test exec time sec: 219.959

----- TestDFSIO ----- : read
           Date & time: Thu Jun 13 14:01:01 WIB 2019
       Number of files: 4
Total MBytes processed: 1024.0
     Throughput mb/sec: 18.00692844707827
Average IO rate mb/sec: 18.323461532592773
 IO rate std deviation: 2.501963465819598
    Test exec time sec: 64.316
我认为,有了更多的节点,工作将更加并行化并提高吞吐量。为什么在添加新节点时写入操作会大大减少?

1 个答案:

答案 0 :(得分:0)

您的数据大小太小。单个系统可以轻松处理1 GB数据。考虑到这是您正在使用的最大尺寸,看到这些结果就不足为奇了。

将这几个数量级提高到100GB-1TB,否则从这种类型的测试中得出性能结果就没有意义。