Friday, February 19, 2010

Triggering post-Elastic MapReduce steps as parameterized jobs in Hudson

Here at Bizo, the combination of Hudson for cron management, Hive for report generation, and Elastic MapReduce for provisioning compute power has greatly simplified our data processing. Periodically and automatically, our Hudson cron instance generates Hive scripts for us and launches them in EC2.

The main inconvenience with this process is that the results of our Hive jobs are left as one or more obscurely named files in S3. These often need some post-processing to put them into a more friendly form. Unfortunately, EMR doesn't have an easy hook for launching these post-processing tasks -- although we could implement them as MapReduce steps, we'd need to write our own workflows, losing the simplicity of using EMR's simple "--hive-script" flag.

Our solution is to use SimpleDB to store some basic metadata about jobs. Using this metadata, a Hudson job periodically checks the EMR API to determine whether tasks have completed. If so, it then triggers other Hudson jobs that are responsible for processing the results.

Here are some tools that make this process work:

  • Simple script to put data into SimpleDB. Our metadata scheme is to use the jobflow ID as item names and the name/parameters of the jobs to trigger as attributes.

  • The Hudson parameterized build feature. It's not really feasible to create a new Hudson job for each individual report that runs, so we pass parameters to Hudson so the post-processing step can figure out where the results are in S3 and what to do with them. It's not well-documented how to do this programatically (as opposed to from the web interface); the solution is to send some JSON to the build url.

  • The Trigger Script. This is the script that periodically runs on our cron server to check if a post-processing step should be triggered. The JSON format for parameterized jobs is described in the comments of this file.



The end result is that a job can run an EMR job and configure a post-processing step for itself with the following commands:


JOB_ID=`elastic-mapreduce --create --hive-script --arg ${s3.location}` | grep "Created job flow" | awk '{ print $4 }' -`

simpledb-put.rb -d ${metadata.domain} -i $JOB_ID "next_on_cron_server_job_name=post-processing-step" "next_on_cron_server_job_params={\"parameter\": [{\"name\":\"PARAM1\", \"value\":\"VALUE1\" }]}" "next_on_cron_server_triggered=false"


This launches the hive script in the specified s3.location and configures the "post-processing-step" job on the cron server to run with the parameter "PARAM1=VALUE1".

Thursday, January 7, 2010

Scala Supports Non-Local Returns

Writing some Scala code today, I found myself using non-local returns without even thinking about it. After realizing what I had done, I dug a little deeper to see what was really going on.

Take this completely made up, nonsensical example:


object Foo {
def main(args: Array[String]) {
foo(List(1, 2, 3))
}

def foo(l: List[Int]): Int = {
l.foreach { (i) =>
println(i)
return 5
}
return 10
}
}


This code will print "1" and then exit.

Perhaps this is obvious, that the "return 5" applies to the "foo" method, so the values 2 and 3 in "l" will not have a chance to be printed.

However, think about what is going on under the covers--Scala is passing the foreach method an anonymous inner class with a "void apply(int i)" method. And inside of that "apply" method is the code between the "{ (i) => ... }".

So, how does code inside of the "apply" method cause its caller to perform an early exit, without the "foreach" even knowing about it?

Exceptions.

Here is the decompiled version of "print":


public int print(List l) {
Object localObject = new Object();
int exceptionResult1 = 0;
try {
l.foreach(new AbstractFunction1() {
public static final long serialVersionUID = 0L;

public final Nothing. apply(int i) {
Predef..MODULE$.println(BoxesRunTime.boxToInteger(i));
// here is the "return"--it puts "5" into an exception
throw new NonLocalReturnException(
this.nonLocalReturnKey1$1,
BoxesRunTime.boxToInteger(5));
}
});
return 10;
} catch (NonLocalReturnException localNonLocalReturnException) {
if (localNonLocalReturnException.key() == localObject) {
// get "5" back out of the exception
return BoxesRunTime.unboxToInt(localNonLocalReturnException.value());
}
throw localNonLocalReturnException;
}
}


I think the decompiler got a little confused with "nonLocalReturnKey", but you can see the basic idea is that any early return inside of a closure is converted into an exception that is then caught outside of the closure where a proper return call can be done.

I personally think this is handy, once you know what is going on. But from what I've picked up, any closures that make it into Java 7 will not support non-local returns and instead disallow the "return" keyword inside of closures. Which, I guess, at this point any Java closures are better than no closures at all.

Tuesday, December 15, 2009

amazon ec2 spot instances

Yesterday Amazon announced EC2 Spot Instances. The idea is that you can bid on unused EC2 instance time. The 'Spot price' is determined periodically by Amazon based on availability and demand for the instances. If your bid is higher than the spot price, you will get an instance and only pay the spot price. Of course, your instance may be terminated at any time, but the nice thing is that unlike the normal ec2 pricing, here you only pay for full hours of usage.

To check out the price history of small linux instances, download the new release of the ec2-api-tools and run:


ec2-describe-spot-price-history --instance-type m1.small -d Linux/UNIX -H


Running this last night, I saw prices that looked like (times PST):



It looks like there's a substantial discount here with prices ranging from $0.025 to $0.035 per hour (the normal ec2 price is $0.085/hr).

Since I'm in the middle of reading How to Cheat at Everything, one of my first thoughts was why not just bid say $0.10/hour? In this way, you're unlikely to get outbid, but you'll probably stand to save significantly for a large part of the day. Now I'm thinking this probably isn't quite a free market... If amazon needs capacity to satisfy reserved instances, or even regular ec2 instances, maybe they'll just kill off these machines to make room.

Still, this is really very cool. A great option for doing a lot of offline batch processing. I hope we start to see support for taking advantage of this type of model in hadoop. It's also exciting to think that one day maybe we'll see something like this across providers -- bid for time across amazon, sun, etc.

Update: Some nice charts by Tim Lossen at cloudexchange.org.

Friday, December 4, 2009

github spam?

I just happened to land on the github recent repositories page, and noticed a ton of spam:



A bunch of different users and projects advertising movie downloads. There's no project content, of course, just a "homepage" that points to a target url...

At first I was thinking, wow, these are some crazy spammers -- using git as a tool for spam! But on closer look, it seems like they're just hitting the website automating account signup and new repository actions.

Still, spam on github? Crazy! I guess no website is safe these days. If you're hosting user generated content, you need to think about detecting and blocking spam and automation of user activity.

Monday, November 30, 2009

quick script: open hadoop jobtracker UI with elastic map reduce

If you've ever logged into the hadoop master with amazon's elastic map reduce, you'll see something like:

The Hadoop UI can be accessed via the command: lynx http://localhost:9100/

Great, but lynx?.. not as nice as firefox or safari...

It's easy enough to do some ssh port forwarding so you can use your browser of choice and access the hadoop UI from your machine.

But, after getting tired of typing in the ssh options a bunch of times, I finally put together a short script that automates it a bit. The script takes in the public hostname of your hadoop master (you can get this from elastic-mapreduce --list), then picks a random port number, sets up the ssh forwarding, and opens the page in a new browser window.

I call it hcon for 'hadoop console'. After configuring the script with the path to your emr key file, you run it like:

hcon ec2-XXX-XXX-XXX-XXX.compute-1.amazonaws.com

Here's the full script, but in case you're curious the magic lines (wrapped) are:

ssh -f -N -o "StrictHostKeyChecking no" \
-L ${LPORT}:localhost:9100 \
-i ${KEYFILE} hadoop@${HOST}
$BROWSER http://localhost:${LPORT}

(Yes, for this, I turn off StrictHostKeyChecking).

Anyway, try it out and let me know if it's helpful at all.

Friday, November 13, 2009

Monday, November 2, 2009

Using Hudson to manage crons

We've been using Hudson for several months now to manage our builds -- we probably have 80-90 different projects that it's responsible for. It's an awesome system for continuous integration and testing.


It's also an awesome system for scheduling and managing generic jobs. We've only just begun to use it as a cron server, but it's clear that it has numerous advantages over the more traditional way of using the unix cron service directly.


  • Notification plugins -- Hudson can be easily configured to send email and Jabber notifications when cron jobs start, succeed, or fail. You can also track your scheduled jobs via RSS.

  • Stdout/Sterr logging -- Hudson saves the stdout and stderr from each run automatically.

  • SCM integration -- if you need to update a job, just check the changes into SVN (or whatever SCM system you use). Hudson will automatically pick up the changes the next time your job is run.

  • Nice web interface -- never underestimate the productivity gains from having a good UI. It can be surprisingly tricky to determine exactly which crons are running on a generic Unix box. Not so with Hudson.


At Bizo, we believe that developers should be getting their hands dirty in the operational aspects of their projects -- Hudson gives us an easy interface for managing our scheduled jobs using the same tools that we're familiar with for managing our build processes. Hudson is such a great tool for continuous integration that it's easy to overlook how good it is at the simpler task of managing generic scheduled jobs.