Hadoop - 在完成reduce任务后,map任务继续

时间:2013-05-20 05:09:46

标签: hadoop

我在大约500个节点的集群上运行Hadoop 1.0.0版。 我的工作有大约3000个地图任务和10个减少任务。 大约4小时后(正如预期的那样)完成地图任务。 然后很快完成reduce任务,结果在我的输出目录中都可用。然而,jobtracker认为某些map任务已经失败并开始重新执行它们。执行和待处理减少任务的数量保持为零。 最终大约8小时后,最后一个映射任务最终成功完成,作业被标记为已成功完成。

任何想法???


以下是一些jobtracker日志文件的摘录:

// map tasks all complete, eg:
2013-05-20 10:50:59,742 INFO org.apache.hadoop.mapred.JobInProgress:   
Task 'attempt_201305131710_0007_m_000430_0' has completed task_201305131710_0007_m_000430 successfully.

//reduce tasks all complete:

2013-05-20 13:38:34,040 INFO org.apache.hadoop.mapred.JobInProgress:         Task 'attempt_201305131710_0007_r_000009_0' has completed task_201305131710_0007_r_000009 successfully.
2013-05-20 13:38:34,142 INFO org.apache.hadoop.mapred.JobInProgress:    
Task 'attempt_201305131710_0007_r_000004_0' has completed task_201305131710_0007_r_000004 successfully.
2013-05-20 13:38:34,204 INFO org.apache.hadoop.mapred.JobInProgress:    
Task 'attempt_201305131710_0007_r_000008_0' has completed task_201305131710_0007_r_000008 successfully.
2013-05-20 13:38:34,745 INFO org.apache.hadoop.mapred.JobInProgress:   
Task 'attempt_201305131710_0007_r_000002_0' has completed task_201305131710_0007_r_000002 successfully.
2013-05-20 13:38:35,521 INFO org.apache.hadoop.mapred.JobInProgress:    
Task 'attempt_201305131710_0007_r_000003_0' has completed task_201305131710_0007_r_000003 successfully.
2013-05-20 13:38:36,196 INFO org.apache.hadoop.mapred.JobInProgress:   
Task 'attempt_201305131710_0007_r_000007_0' has completed task_201305131710_0007_r_000007 successfully.
2013-05-20 13:38:36,276 INFO org.apache.hadoop.mapred.JobTracker: Adding tracker tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:1295 to host HN301-1657.labs.edu.au
2013-05-20 13:38:36,469 INFO org.apache.hadoop.mapred.JobInProgress:   
Task 'attempt_201305131710_0007_r_000005_0' has completed task_201305131710_0007_r_000005 successfully.
2013-05-20 13:38:36,598 INFO org.apache.hadoop.mapred.JobInProgress:  
Task 'attempt_201305131710_0007_r_000006_0' has completed task_201305131710_0007_r_000006 successfully.
2013-05-20 13:38:36,612 INFO org.apache.hadoop.mapred.JobInProgress:   
Task 'attempt_201305131710_0007_r_000000_0' has completed task_201305131710_0007_r_000000 successfully.
2013-05-20 13:38:40,388 INFO org.apache.hadoop.mapred.JobInProgress:  
Task 'attempt_201305131710_0007_r_000001_0' has completed task_201305131710_0007_r_000001 successfully.
2013-05-20 13:44:12,795 INFO org.apache.hadoop.mapred.JobTracker: Lost tracker 'tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896'

//As the reduce tasks are reporting success, the job tracker detects that one of the job trackers has died and so restarts it.
//Each of the jobs previously completed successfully by that task tracker are then reexecuted
2013-05-20 13:44:12,795 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_000430_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,795 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_000571_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,796 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_001612_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,796 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_001629_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,796 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_001892_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,796 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_002424_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,796 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_002437_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,796 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_002696_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,796 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_003130_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,796 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_003149_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,796 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_003187_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_003275_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_003358_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_003437_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_003451_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_003478_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0007_m_003506_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.TaskInProgress: Error from attempt_201305131710_0010_m_000021_0: Lost task tracker: tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0010_m_000021_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_000430_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_000571_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_001612_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_001629_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_001892_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_002424_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_002437_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_002696_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_003130_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_003149_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_003187_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_003275_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_003358_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_003437_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_003451_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_003478_0'
2013-05-20 13:44:12,797 INFO org.apache.hadoop.mapred.JobTracker: Removing task 'attempt_201305131710_0007_m_003506_0'
2013-05-20 13:44:12,917 INFO org.apache.hadoop.mapred.JobTracker: Adding task (TASK_CLEANUP) 'attempt_201305131710_0010_m_000021_0' to tip task_201305131710_0010_m_000021, for   tracker 'tracker_HN301-1654.labs.edu.au:127.0.0.1/127.0.0.1:1100'
2013-05-20 13:44:13,760 INFO org.apache.hadoop.mapred.JobInProgress: Choosing a failed task task_201305131710_0007_m_000430
2013-05-20 13:44:13,761 INFO org.apache.hadoop.mapred.JobTracker: Adding task (MAP) 'attempt_201305131710_0007_m_000430_1' to tip task_201305131710_0007_m_000430, for tracker 'tracker_ZC329-0001.labs.edu.au:127.0.0.1/127.0.0.1:1113'

1 个答案:

答案 0 :(得分:0)

您可能想要检查每个群集节点的配置:

tracker_HN301-1657.labs.edu.au:127.0.0.1/127.0.0.1:3896

jobTracker正试图通过回送地址联系TaskTracker节点。检查每个节点/ etc / hosts文件的内容以检查它们是否正确(并且最好了解群集中的每个其他节点,以便您可以避免DNS查找成本。)

我不是说这是你的问题的原因,但它肯定是不对的,应该是你追踪的东西