最近需要把基于Hadoop的MapReduce程序集成到一个大的用C/C++编写的框架中,需要在make的时候自动将MapReduce应用进行编译和打包。这里以简单的WordCount1为例说明具体的实现细节,注意:hadoop版本为2.4.0.
源代码包含两个文件,一个是WordCount1.java是具体的对单词计数实现的逻辑;第二个是CounterThread.java,其中简单的当前处理的行数做一个统计和打印。代码分别见附1. 编写makefile的关键是将hadoop提供的jar包的路径全部加载进来,看到网上很多资料都自己实现一个脚本把hadoop目录下所有的.jar文件放到一个路径中,然后进行编译,这种做法太麻烦了。当然也有些简单的办法,但是都是比较老的hadoop版本如0.20之类的。
其实,hadoop提供了一个命令hadoop classpath可以获得包含所有jar包的路径.所以只需要用 javac -classpath "`hadoop classpath`" *.java 便可,然后使用jar -cvf对class文件进行打包就可以了。
--------------------------------------分割线 --------------------------------------
Ubuntu 13.04上搭建Hadoop环境
Ubuntu 12.10 +Hadoop 1.2.1版本集群配置
--------------------------------------分割线 --------------------------------------
具体的Makefile代码如下:
SRC_DIR = src/mypackage/*.java
CLASS_DIR = bin
TARGET_JAR = WordCount
all:$(TARGET_JAR)
$(TARGET_JAR): $(SRC_DIR)
mkdir -p $(CLASS_DIR)
# javac -classpath `$(HADOOP) classpath` -d $(CLASS_DIR) $(SRC_DIR)
javac -classpath "`hadoop classpath`" src/mypackage/*.java -d $(CLASS_DIR) -Xlint
jar -cvf $(TARGET_JAR).jar -C $(CLASS_DIR) ./
clean:
rm -rf $(CLASS_DIR) *.jar
make一下:
lichao@ubuntu:WordCount1$ make
mkdir -p bin
javac -classpath "`hadoop classpath`" src/mypackage/*.java -d bin -Xlint
warning: [path] bad path element "/home/lichao/Software/hadoop/hadoop-src/hadoop-2.4.0-src/hadoop-dist/target/hadoop-2.4.0/share/hadoop/common/lib/jaxb-api.jar": no such file or directory
warning: [path] bad path element "/home/lichao/Software/hadoop/hadoop-src/hadoop-2.4.0-src/hadoop-dist/target/hadoop-2.4.0/share/hadoop/common/lib/activation.jar": no such file or directory
warning: [path] bad path element "/home/lichao/Software/hadoop/hadoop-src/hadoop-2.4.0-src/hadoop-dist/target/hadoop-2.4.0/share/hadoop/common/lib/jsr173_1.0_api.jar": no such file or directory
warning: [path] bad path element "/home/lichao/Software/hadoop/hadoop-src/hadoop-2.4.0-src/hadoop-dist/target/hadoop-2.4.0/share/hadoop/common/lib/jaxb1-impl.jar": no such file or directory
warning: [path] bad path element "/home/lichao/Software/hadoop/hadoop-src/hadoop-2.4.0-src/hadoop-dist/target/hadoop-2.4.0/share/hadoop/yarn/lib/jaxb-api.jar": no such file or directory
warning: [path] bad path element "/home/lichao/Software/hadoop/hadoop-src/hadoop-2.4.0-src/hadoop-dist/target/hadoop-2.4.0/share/hadoop/yarn/lib/activation.jar": no such file or directory
warning: [path] bad path element "/home/lichao/Software/hadoop/hadoop-src/hadoop-2.4.0-src/hadoop-dist/target/hadoop-2.4.0/share/hadoop/yarn/lib/jsr173_1.0_api.jar": no such file or directory
warning: [path] bad path element "/home/lichao/Software/hadoop/hadoop-src/hadoop-2.4.0-src/hadoop-dist/target/hadoop-2.4.0/share/hadoop/yarn/lib/jaxb1-impl.jar": no such file or directory
warning: [path] bad path element "/home/lichao/Software/hadoop/hadoop-src/hadoop-2.4.0-src/hadoop-dist/target/hadoop-2.4.0/contrib/capacity-scheduler/*.jar": no such file or directory
src/mypackage/WordCount1.java:61: warning: [deprecation] Job(Configuration,String) in Job has been deprecated
Job job = new Job(conf, "WordCount1"); //建立新job
^
10 warnings
jar -cvf WordCount.jar -C bin ./
added manifest
adding: mypackage/(in = 0) (out= 0)(stored 0%)
adding: mypackage/WordCount1.class(in = 1970) (out= 1037)(deflated 47%)
adding: mypackage/CounterThread.class(in = 1760) (out= 914)(deflated 48%)
adding: mypackage/WordCount1$IntSumReducer.class(in = 1762) (out= 749)(deflated 57%)
adding: mypackage/WordCount1$TokenizerMapper.class(in = 1759) (out= 762)(deflated 56%)
adding: log4j.properties(in = 476) (out= 172)(deflated 63%)
虽然有warning,但是不影响结果。编译后,我们来简单的测试一下。