Pokazywanie postów oznaczonych etykietą MapReduce. Pokaż wszystkie posty
Pokazywanie postów oznaczonych etykietą MapReduce. Pokaż wszystkie posty

poniedziałek, 25 lutego 2013

maven-hadoop-plugin

Submitting hadoop job on remote machine is not a complicated process but it takes a lot time, it could be 10 minutes or sometime 15. There are a lot steps to do to get the final result of map reduce and download it to local file system.

-compile map reduce code and build jar file
-upload jar to remote server via ftp
-connect to server via ssh client
-prepare input data in HDFS (usually only once)
-submit job using ‘hadoop jar…’ command
-copy map reduce output from HDFS to remote server local file system
-download output to local machine

Like I mentioned, doing those entire steps manually takes time. I was seeking for some ways to do it faster and easier. After short research, because I didn’t find any tool or solution, so I state the easier and the best way is to develop something myself. I was wondering how to do this. I could code some standalone application for doing this but it would not be enough comfortable again and require few steps from the user like switching from java IDE to another window and finding jar in file system. I thought maybe it will be better to write eclipse plugin. Everything would be in one place but would have some weakness also – no usage outside of eclipse.  Next thought was Maven. Integrated with … everything could be used in the console and what important plugin development for it is easy and pleasant, so I took this idea started development.

The result of my work looks really well. Now I submit my job by ‘one button click’. I have my maven execution configured in eclipse (hadoop:execute)



 My sample wordcount need to have maven-hadoop-plugin configured in its pom.xmll file

<build>
 <plugins>
  <plugin>
   <groupId>org.apache.maven.plugins</groupId>
   <artifactId>maven-hadoop-plugin</artifactId>
   <version>0.0.1-SNAPSHOT</version>
   <configuration>
    <host>ipAddress</host>
    <login>login</login>
    <password>password</password>
    <outputDir>output30</outputDir>
    <hdfsOutputDir>
      /books/users/gkolpu/output30
    </hdfsOutputDir>
    <hdfsInputDir>
      /books/input
    </hdfsInputDir>
    <className>
      org.gkolpu.hadoop.BookWordCounter
    </className>
    <jarName>WordCounter.jar</jarName>
   </configuration>
  </plugin>
 </plugins>
</build>



Running my maven hadoop:execute goal from eclipse I see all logs coloured in IDE console



When the job finished on remote server map reduce output is downloaded automatically to target directory




Output files can be easily opened in IDE editors




You can find this plugin with source code on my git-hub repository. There is one only one goal but this project is still under development and other maven goals are planned.

https://github.com/gkolpuc/maven-hadoop-plugin

piątek, 4 stycznia 2013

Map Reduce implementation with Hadoop

Hadoop is an open source framework which supports big data distributed application. One od main features is MapReduce algorithm implementation. Hadoop gives us opportunity to use it's API to implement our own Maping and Reducing. To run your first hadoop job you will need to implement generic Mapper and Reducer classes. See examples below.
public class Map extends MapReduceBase implements
  Mapper {

 private final IntWritable one = new IntWritable(1);
 private Text word = new Text();

 public void map(LongWritable key, Text value,
   OutputCollector output, Reporter reporter)
   throws IOException {

  String line = value.toString();
  StringTokenizer tokenizer = new StringTokenizer(line);
  while (tokenizer.hasMoreTokens()) {
   word.set(tokenizer.nextToken());
   output.collect(word, one);
  }
 }
}
public class Reduce extends MapReduceBase implements
  Reducer {
 public void reduce(Text key, Iterator values,
   OutputCollector output, Reporter reporter)
   throws IOException {
  int sum = 0;

  while (values.hasNext()) {
   sum += values.next().get();
  }

  output.collect(key, new IntWritable(sum));
 }
}
Last thing you need to do is clip them together using Job Configuration
public class Job {

 public static final void main(String[] args) throws Exception {

  JobConf conf = new JobConf(Job.class);
  conf.setJobName("Hadoop-Workshop-GKOLPU");

  conf.setOutputKeyClass(Text.class);
  conf.setOutputValueClass(IntWritable.class);

  conf.setMapperClass(Map.class);
  conf.setCombinerClass(Reduce.class);
  conf.setReducerClass(Reduce.class);

  conf.setInputFormat(TextInputFormat.class);
  conf.setOutputFormat(TextOutputFormat.class);

 FileInputFormat.setInputPaths(conf,new Path(args[1]));
 FileOutputFormat.setOutputPath(conf,new Path(args[2]));

  JobClient.runJob(conf);
 }
}