1.编译源码

  1. 环境准备
  2. maven(下载安装,配置环境变量,修改sitting.xml加阿里云镜像)
  3. gcc-c++
  4. zlib-devel
  5. autoconf
  6. automake
  7. libtool
  8. 通过yum安装即可,
  9. yum -y install gcc-c++ lzo-devel zlib-devel autoconf automake libtool
  10. 1. 下载、安装并编译LZO
  11. wget http://www.oberhumer.com/opensource/lzo/download/lzo-2.10.tar.gz
  12. tar -zxvf lzo-2.10.tar.gz
  13. cd lzo-2.10
  14. ./configure -prefix=/usr/local/hadoop/lzo/
  15. make
  16. make install
  17. 2. 编译hadoop-lzo源码
  18. 下载hadoop-lzo的源码,下载地址:
  19. https://github.com/twitter/hadoop-lzo/archive/master.zip
  20. 解压之后,修改pom.xml
  21. <hadoop.current.version>2.7.2</hadoop.current.version>
  22. 声明两个临时环境变量
  23. export C_INCLUDE_PATH=/usr/local/hadoop/lzo/include
  24. export LIBRARY_PATH=/usr/local/hadoop/lzo/lib
  25. export CFLAGS=-m64
  26. export CXXFLAGS=-m64
  27. 编译
  28. cd hadoop-lzo-master
  29. 执行maven编译命令
  30. mvn package -Dmaven.test.skip=true
  31. 进入targethadoop-lzo-0.4.21-SNAPSHOT.jar 即编译成功的hadoop-lzo组件

2.配置

  1. 将编译好后的hadoop-lzo-0.4.20.jar 放入hadoop-2.7.2/share/hadoop/common/
  2. xsync hadoop-lzo-0.4.20.jar 并完成集群的分发
  3. core-site.xml增加配置支持LZO压缩
  4. <property>
  5. <name>io.compression.codecs</name>
  6. <value>
  7. org.apache.hadoop.io.compress.GzipCodec,
  8. org.apache.hadoop.io.compress.DefaultCodec,
  9. org.apache.hadoop.io.compress.BZip2Codec,
  10. org.apache.hadoop.io.compress.SnappyCodec,
  11. com.hadoop.compression.lzo.LzoCodec,
  12. com.hadoop.compression.lzo.LzopCodec
  13. </value>
  14. </property>
  15. <property>
  16. <name>io.compression.codec.lzo.class</name>
  17. <value>com.hadoop.compression.lzo.LzoCodec</value>
  18. </property>
  19. 配置文件集群文件分发
  20. xsync core-site.xml
  21. 重启集群

3.测试

  1. yarn jar /export/server/hadoop2.7/share/hadoop/mapreduce/hadoop-mapreduce-examples-2.7.5.jar wordcount -Dmapreduce.output.fileoutputformat.compress=true -Dmapreduce.output.fileoutputformat.compress.codec=com.hadoop.compression.lzo.LzopCodec /data/mr/wordcount/input /data/mr/wordcount/output
  2. lzo文件创建索引
  3. hadoop jar ./share/hadoop/common/hadoop-lzo-0.4.20.jar com.hadoop.compression.lzo.DistributedLzoIndexer /output