在oozie工作流程中使用hcatalog进行sqoop操作有问题

时间:2018-12-29 07:12:06

标签: hive sqoop oozie hcatalog oozie-workflow

当我使用sqoop export命令将数据从蜂巢导出到mirosoft sql server时,在ambary-views中将sqoop actin与hcatalog一起使用时遇到问题。

以下命令在shell中正确运行,效果很好。

sqoop export --connect 'jdbc:sqlserver://x.x.x.x:1433;useNTLMv2=true;databasename=BigDataDB'  --connection-manager org.apache.sqoop.manager.SQLServerManager --username 'DataApp' --password 'D@t@User' --table tr1 --hcatalog-database temporary --catalog-table 'daily_tr'

但是当我在oozie工作流程中使用此命令创建sqoop操作时,出现以下错误:

Failing Oozie Launcher, Main class [org.apache.oozie.action.hadoop.SqoopMain], main() threw exception, org/apache/hive/hcatalog/mapreduce/HCatOutputFormat
java.lang.NoClassDefFoundError: org/apache/hive/hcatalog/mapreduce/HCatOutputFormat
        at org.apache.sqoop.mapreduce.ExportJobBase.runExport(ExportJobBase.java:432)
        at org.apache.sqoop.manager.SQLServerManager.exportTable(SQLServerManager.java:192)
        at org.apache.sqoop.tool.ExportTool.exportTable(ExportTool.java:81)
        at org.apache.sqoop.tool.ExportTool.run(ExportTool.java:100)
        at org.apache.sqoop.Sqoop.run(Sqoop.java:147)
        at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:76)
        at org.apache.sqoop.Sqoop.runSqoop(Sqoop.java:183)
        at org.apache.sqoop.Sqoop.runTool(Sqoop.java:225)
        at org.apache.sqoop.Sqoop.runTool(Sqoop.java:234)
        at org.apache.sqoop.Sqoop.main(Sqoop.java:243)
        at org.apache.oozie.action.hadoop.SqoopMain.runSqoopJob(SqoopMain.java:171)
        at org.apache.oozie.action.hadoop.SqoopMain.run(SqoopMain.java:153)
        at org.apache.oozie.action.hadoop.LauncherMain.run(LauncherMain.java:75)
        at org.apache.oozie.action.hadoop.SqoopMain.main(SqoopMain.java:50)
        at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
        at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)
        at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
        at java.lang.reflect.Method.invoke(Method.java:498)
        at org.apache.oozie.action.hadoop.LauncherMapper.map(LauncherMapper.java:231)
        at org.apache.hadoop.mapred.MapRunner.run(MapRunner.java:54)
        at org.apache.hadoop.mapred.MapTask.runOldMapper(MapTask.java:453)
        at org.apache.hadoop.mapred.MapTask.run(MapTask.java:343)
        at org.apache.hadoop.mapred.YarnChild$2.run(YarnChild.java:170)
        at java.security.AccessController.doPrivileged(Native Method)
        at javax.security.auth.Subject.doAs(Subject.java:422)
        at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1869)
        at org.apache.hadoop.mapred.YarnChild.main(YarnChild.java:164)
Caused by: java.lang.ClassNotFoundException: org.apache.hive.hcatalog.mapreduce.HCatOutputFormat
        at java.net.URLClassLoader.findClass(URLClassLoader.java:381)
        at java.lang.ClassLoader.loadClass(ClassLoader.java:424)
        at sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:338)
        at java.lang.ClassLoader.loadClass(ClassLoader.java:357)
        ... 27 more

要解决此错误,请执行以下操作:

  • 在workflow.xml所在的文件夹下,我创建文件夹lib,并将所有来自sharedlibDir(/ user / oozie / share / lib / lib_lib_201806281525405 / hive的hive jar文件放在其中。

我的目标是这样做,组件可以识别hcatalog jar文件和类路径,因此我不确定,也许我不应该这样做,并针对此错误采取不同的解决方案

无论如何,该错误已更改如下:

Failing Oozie Launcher, Main class [org.apache.oozie.action.hadoop.SqoopMain], main() threw exception, org.apache.hadoop.hive.shims.HadoopShims.g
etUGIForConf(Lorg/apache/hadoop/conf/Configuration;)Lorg/apache/hadoop/security/UserGroupInformation;
java.lang.NoSuchMethodError: org.apache.hadoop.hive.shims.HadoopShims.getUGIForConf(Lorg/apache/hadoop/conf/Configuration;)Lorg/apache/hadoop/sec
urity/UserGroupInformation;
        at org.apache.hive.hcatalog.common.HiveClientCache$HiveClientCacheKey.<init>(HiveClientCache.java:201)
        at org.apache.hive.hcatalog.common.HiveClientCache$HiveClientCacheKey.fromHiveConf(HiveClientCache.java:207)
        at org.apache.hive.hcatalog.common.HiveClientCache.get(HiveClientCache.java:138)
        at org.apache.hive.hcatalog.common.HCatUtil.getHiveClient(HCatUtil.java:564)
        at org.apache.hive.hcatalog.mapreduce.InitializeInput.getInputJobInfo(InitializeInput.java:104)
        at org.apache.hive.hcatalog.mapreduce.InitializeInput.setInput(InitializeInput.java:86)
        at org.apache.hive.hcatalog.mapreduce.HCatInputFormat.setInput(HCatInputFormat.java:85)
        at org.apache.hive.hcatalog.mapreduce.HCatInputFormat.setInput(HCatInputFormat.java:63)
        at org.apache.sqoop.mapreduce.hcat.SqoopHCatUtilities.configureHCat(SqoopHCatUtilities.java:349)
        at org.apache.sqoop.mapreduce.ExportJobBase.runExport(ExportJobBase.java:433)
        at org.apache.sqoop.manager.SQLServerManager.exportTable(SQLServerManager.java:192)
        at org.apache.sqoop.tool.ExportTool.exportTable(ExportTool.java:81)
        at org.apache.sqoop.tool.ExportTool.run(ExportTool.java:100)
        at org.apache.sqoop.Sqoop.run(Sqoop.java:147)
        at org.apache.hadoop.util.ToolRunner.run(ToolRunner.java:76)
        at org.apache.sqoop.Sqoop.runSqoop(Sqoop.java:183)
        at org.apache.sqoop.Sqoop.runTool(Sqoop.java:225)
        at org.apache.sqoop.Sqoop.runTool(Sqoop.java:234)
        at org.apache.sqoop.Sqoop.main(Sqoop.java:243)
        at org.apache.oozie.action.hadoop.SqoopMain.runSqoopJob(SqoopMain.java:171)
        at org.apache.oozie.action.hadoop.SqoopMain.run(SqoopMain.java:153)
        at org.apache.oozie.action.hadoop.LauncherMain.run(LauncherMain.java:75)
        at org.apache.oozie.action.hadoop.SqoopMain.main(SqoopMain.java:50)
        at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
        at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)
        at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
        at java.lang.reflect.Method.invoke(Method.java:498)
        at org.apache.oozie.action.hadoop.LauncherMapper.map(LauncherMapper.java:231)
        at org.apache.hadoop.mapred.MapRunner.run(MapRunner.java:54)
        at org.apache.hadoop.mapred.MapTask.runOldMapper(MapTask.java:453)
        at org.apache.hadoop.mapred.MapTask.run(MapTask.java:343)
        at org.apache.hadoop.mapred.YarnChild$2.run(YarnChild.java:170)
        at java.security.AccessController.doPrivileged(Native Method)

版本:

HDP 2.6.5.0

纱线2.7.3

配置单元1.2.1000

sqoop 1.4.6

oozie 4.2.0

请帮助我解决错误和问题,为什么sqoop命令在shell中正确运行,但是在oozie工作流程中却有错误?

2 个答案:

答案 0 :(得分:0)

我不知道这是否是罪魁祸首。我一年前在HDP中使用Sqoop 1.4.x遇到了此问题,它显示死于一些不相关的失败原因。

从命令行在sqoop命令下运行时,它将成功运行。

sqoop export --connect 'jdbc:sqlserver://x.x.x.x:1433;useNTLMv2=true;databasename=BigDataDB'  --connection-manager org.apache.sqoop.manager.SQLServerManager --username 'DataApp' --password 'D@t@User' --table tr1 --hcatalog-database temporary --catalog-table 'daily_tr'

但是,当您通过Oozie Sqoop操作运行相同的命令时,它不应像下面那样使用单引号(')。

<command>export --connect jdbc:sqlserver://x.x.x.x:1433;useNTLMv2=true;databasename=BigDataDB  --connection-manager org.apache.sqoop.manager.SQLServerManager --username DataApp --password D@t@User --table tr1 --hcatalog-database temporary --catalog-table daily_tr</command>

答案 1 :(得分:0)

我通过以下方法解决了我的问题:

1-在工作流.xml的命令标记中使用(--hcatalog-home / usr / hdp / current / hive-webhcat):

<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<workflow-app xmlns="uri:oozie:workflow:0.5" name="loadtosql">
    <start to="sqoop_export"/>
    <action name="sqoop_export">
        <sqoop xmlns="uri:oozie:sqoop-action:0.4">
            <job-tracker>${resourceManager}</job-tracker>
            <name-node>${nameNode}</name-node>
            <command>export --connect jdbc:sqlserver://x.x.x.x:1433;useNTLMv2=true;databasename=BigDataDB --connection-manager org.apache.sqoop.manager.SQLServerManager --username DataApp--password D@t@User --table tr1 --hcatalog-home /usr/hdp/current/hive-webhcat --hcatalog-database temporary --hcatalog-table daily_tr </command>
            <file>/user/ambari-qa/test/lib/hive-site.xml</file>
            <file>/user/ambari-qa/test/lib/tez-site.xml</file>
        </sqoop>
        <ok to="end"/>
        <error to="kill"/>
    </action>
    <kill name="kill">
        <message>${wf:errorMessage(wf:lastErrorNode())}</message>
    </kill>
    <end name="end"/>
</workflow-app>

2-在hdfs上,在workflow.xml旁边创建lib文件夹,然后将hive-site.xml和tez-site.xml放在其中(从/etc/hive/2.6.5.0-292/0上传hive-site.xml /和tez-site.xml从/etc/tez/2.6.5.0-292/0/到hdfs上的lib文件夹)

根据上面的工作流定义两个文件(hive-site.xml和tez-site.xml)

<file>/user/ambari-qa/test/lib/hive-site.xml</file>
<file>/user/ambari-qa/test/lib/tez-site.xml</file>

3-在job.properties文件中定义以下属性:

oozie.action.sharelib.for.sqoop=sqoop,hive,hcatalog

4-确保/ etc / oozie / conf下的oozie-site.xml具有指定的以下属性。

<property> 
    <name>oozie.credentials.credentialclasses</name>
    <value>hcat=org.apache.oozie.action.hadoop.HCatCredentials</value>
</property>