Thursday, April 14, 2011

Hive Unit Testing

Introduction

Hive has become an extremely important component in our overall software stack. We have numerous ‘mission-critical’ reports that are generated using Hive and want to make sure we can apply our testing processes to Hive scripts in the same way that we apply them to other code artifacts.

A few weeks ago, I was tasked with finding an approach for unit testing our Hive scripts. To my surprise, a Google search for ‘Hive Unit Testing’ yielded relatively few useful results.

I wanted a solution that would allow us to test locally (vs. a solution that would require EMR). Where possible, I prefer local testing because it’s simpler, provides more immediate feedback, and doesn’t require a network.

After reading this post, you will (hopefully) know how to run Hive unit tests in your own environment.

The Approach

After performing some research, I decided on an approach that is part of the Hive project itself.  At a high level, the solution works in the following way:
  • Start up an instance of the Hive CLI
  • Execute a Hive script (positive or negative case)
  • Compare the output (from the CLI) of the script compared to an expected output file
  • Rinse and repeat
The rest of this post discusses the specific steps required to get this solution running in your own environment.

Set up Hive Locally

The first step is to create some Ant tasks for setting up Hive locally. Here’s a snippet of Ant that shows how to do this:


You should now be able to execute ‘ant hive.init’ and have Hive available in the tools directory.

Generate test cases

The developer is responsible for providing the .q Hive files that represent the test cases. There is a code generation step that will create JUnit classes (one for positive test cases, one for negative test cases) given a set of .q files. The Ant snippet below shows how to generate the test classes:


Here are some notes about the key variables above:
  • hive.test.template.dir - the directory where the velocity templates are located for the code generation step.
  • target.hive.positive.query.dir - the directory where positive test cases are located.
  • target.hive.negative.query.dir - the directory where negative test cases are located.
  • hive.positive.results.dir - the directory where expected positive test results are located. The name of this file must be the name of the query file appened by ‘.out’. For example, if the test query file is named hive_test.q then the results file must be named hive_test.q.out.
  • hive.negative.results.dir - the directory where expected negative test results are located.
  • qfile - This variable should be specified if you want to generate a test class with a single test case. For example, if you have a test file named hive_test.q, then you would set the value of this property to hive_test (e.g. ant -Dqfile=hive_test hive.gen.test).
  • qfile_regex - Similar in functionality to qfile, this variable should be set to a regular expression that will match the test files that you want to generate tests for.
The test classes are generated from velocity template files. You can find examples of the template from the Hive codebase here: https://github.com/apache/hive/blob/trunk/ql/src/test/templates/TestCliDriver.vm
https://github.com/apache/hive/blob/trunk/ql/src/test/templates/TestNegativeCliDriver.vm

The above files can basically be used as-is, but you will need to provide your own Test Helper class, QTestUtil, and update its package location accordingly in the templates.

QTestUtil

QTestUtil contains code for:
  • starting up hive
  • executing a query file
  • comparing the results to expected results
  • running cleanup between tests
  • shutting down hive
You can find the one from the Hive project here: https://github.com/apache/hive/blob/trunk/ql/src/test/org/apache/hadoop/hive/ql/QTestUtil.java

The main modifications you will want to make to this file are deletions as there is some Hive project specific set up code that you will not need in your environment.

Executing the tests

After you have generated the tests, you can execute them by creating a target with the junit task. Here is some sample Ant for doing this:



Conclusion

This post outlined a solution for unit testing Hive scripts. Another nice aspect of this approach that I failed to mention is that it’s based on JUnit so you can use your existing code coverage tools with it (we use Cobertura) to get coverage information when testing custom UDFs. Also, I should mention that I used Hive 0.6.0 when putting this together.

Wednesday, April 6, 2011

Hive 0.7 no longer auto-downloads transform scripts

I ran into a bit of a surprise moving a Hive 0.5 script to Hive 0.7 the other day.

Previously, in Hive 0.5, we called our Java transform code like:

insert overwrite table the_table
select
transform(...)
using 'java -cp s3://bucket-name/code.jar MapperClassName'

Behind the scenes, before actually calling the "java" executable, Hive would inspect each of the arguments and, if it found an "s3://..." URL, download that file from S3 to a local copy, and then pass the path to the local copy to your program.

This was convenient as then your external "java" executable didn't have to know anything about S3, how to authenticate with it, etc.

However, in Hive 0.7, this no longer works. Perhaps for the understandable reason that if you did want to pass the literal string "s3://..." to your mapper class, Hive implicitly interjecting on your behalf may not be what you want, and, AFAIK, you had no way to avoid it.

So, now an explicit "add file" command is required, e.g.:

add file s3://bucket-name/code.jar
insert overwrite table the_table
select
transform(...)
using 'java -cp code.jar MapperClassName'

The add file command downloads code.jar to the local execution directory (without any bucket name/path mangling like in Hive 0.5), and then your transform arguments can reference the local file directly.

All in all, a pretty easy fix, but rather frustrating to figure out given the long cycle time of EMR jobs.

Also, kudos to this post in the AWS developer forums that describes the same problem and solution:


Friday, March 11, 2011

On Building a Kick Ass Engineering Team -- Part 1


We started Bizo just about three years ago with the goal of building a great business and a world-class engineering team. One of the first things I did was write down what I thought it would take to create a kick-ass engineering team.

I figured it would be good to share this list with others and discuss why I thought the items would help build a great engineering team.

Here is what I wrote:

Let's review these in a bit more detail...

Commitment to Discipline
"Commitment to discipline" is perhaps the most important characteristic for any super productive team. Software development is full of moments that try to pull you away from the task at hand or the problem that you should be trying to solve. Those weak moments where you get pulled into refactoring some sub-system end up being a huge time sink. Being able to recognize that while it would be great to refactor foo, it is not actually a requirement of "getting shit done" and can wait.

Must be cultural
I believe that the culture of any startup comes from the founders and early employees and more generally the culture of any business comes from the leaders. Therefore, it is absolutely imperative for these leaders to define what they want the culture to be and furthermore act on that definition. Culture is self-reinforcing: actions create culture which creates actions which strengthen culture. Just like a habit, once the culture is formed it is difficult to change it so it pays to be purposeful when creating your company's culture. You will have to live with it good or bad.

The 3 Cs
Communication, Communication Communication! (I stole this directly from my high school baseball coach Mr. Barden.) A major part of engineering success comes down to communication. What should we build, how should we build it, how can I plug into your system, how does this code work, etc. We spend an enormous amount of time communicating and it shows. We ensure communication through design reviews for all features/projects, and code reviews for every single line of code that makes it into production. Code reviews are an amazing place to learn, ensure quality and share knowledge. We even spend time communicating about non-Bizo related stuff that we find interesting through "Lunch and Learns".

I'll save how we go about managing all this communication for another post but the ramifications of our commitment to designs reviews, code reviews, and the like result in a team of engineers that are eager to get feedback, humble about their skills and strive to provide objective feedback to others. Objectivity is an extremely valuable characteristic of a great engineer and something that we look for in hiring and something that we strive to develop and promote.

Testing
Testing is a huge part of what we believe in here at Bizo. Beyond the obviousness of ensuring that your code works, testing has some great side effects. For one, it is an investment in future development making it easier to develop features on top of an existing product base and making it easier to "pivot" (I used the word "turn" 3 years ago before pivot was in wide use). Being a big believer in having some down time, I also consider testing an investment in one's weekend because things always seem to go wrong when you are not at work. :) Finally, extensive testing ensures a low amount of code debt which should be any dev team's goal. You always have to pay it off and upfront payments are the cheapest...

Visibility
This should really be part of the 3Cs and I would expand on this even more today. Not only is it important for code metrics, build status and task management to be communicated widely it is extremely important for the efforts of engineers to be communicated. As a company, we have a business standup and an engineering standup where group managers can engage with the teams and get visibility into what is being done. We are also experimenting with bringing engineering earlier into the product development process at the scoping level.

Ownership
Ownership is key to getting the most out of any team including engineering. I think productivity and motivation is a function of happiness and a lot of happiness comes from being trusted to do your job. At a certain point this comes back to hiring. If you can hire people who are great cultural fits and great engineering fits than you must trust them to do their jobs. That trust will make them better employees.

In Summary
To wrap this all up, I think we've done well staying true to what I outlined three years ago and I am extremely happy with the results. I've been around a lot of engineering teams and I've never been more happy (or proud) of what we've been able to accomplish both with the products and code we've shipped and with the culture of engineering we've built.

To Be Continued
In a follow up post I'll discuss what (if anything) I would add or remove from this list if I had to start over today and what I think we should focus on as an engineering team for the next few years.

Thursday, February 24, 2011

"dynamic" columns in Hive

One of the presentations at the HBase meetup the other night was on building a query language on top of HBase. No less than 3 people asked "Why not use Hive?". The main reason given was that Hive is too slow for doing simple selects. But, the other thing they really liked about using HBase was that your columns were dynamic -- it's easy to add new fields to your data.

Most of the data we log is in a simple log file format:
  • One record per line, separated by newline.
  • Each record can have one or more fields. Fields are separated by ^A (\001).
  • Each field is a key/value pair separated by ^B (\002). Field order is not specified.
In practice this bascially looks like:
ts=1298598378404/code=403/message=bad referrer: bizo.com/...
(Where / is ^A and = is ^B).

This is a format we chose pretty early on, way before we ever looked at Hive. It turns out to be a great format:

  • Human readable.
  • Trivial to parse in any language.
  • Dynamic -- easy to add/remove fields from your data.
It also turns out that it works really well with Hive. Our typical Hive table looks something like:

create external table api_logs(d map<string,string>)
partitioned by (...)
row format delimited
fields terminated by '\004'
collection items terminated by '\001'
map keys terminated by '\002'
stored as textfile
;
That is, each row is just a single column, which is a map. At first this seemed a little degenerate to me, but it actually models our data perfectly. There are no guarantees about which fields are available, and it's easy to add/remove fields in the data over time. I should mention that this is really just for our report input, typically our report output will be in a fixed format.

If you're using Hive 0.6 or greater, with Hive View support it's also easy to get the best of both worlds.

create view api_errors(ts, code, message) as
select d["ts"], d["code"], d["message"]
from api_logs
where d["code"] >= 400
;
You can even change the type information or transform the data by including a cast or a UDF as part of your view. Creating a view doesn't cause anything to run, or create any additional storage. Its query conditions are basically just merged with subsequent queries on that view.

Thursday, January 27, 2011

Adventures In GWT-land Part #1: Awkward Baby Steps

Over the past several months we’ve been working on a (super) secret shiny new GWT application (it’s basically going to rock your socks off). This was my first GWT application and coming from a non-Java, non-GWT background where I was used to writing raw Javascript pretty often - it’s been interesting to say the least. What follows is the first in a multi-part series where I’d like to reflect on life in GWT-land and hopefully provide a few cool tips and code samples along the way.

First Steps:

Stepping into GWT for the first time when your used to pure Javascript - or “normal” CSS/HTML front-end development in general is pretty awkward. The biggest adjustment is that all client-side logic is defined in Java classes which later get compiled into several permutations of obfuscated Javascript. GWT also tries to be helpful and obfuscates all of your CSS classes in an attempt to prevent namespace collisions (more on this in a future post). If your thinking that FireBug becomes much less useful when working with GWT your absolutely right...but thanks to some really good debugging tools in GWT, this isn’t a terrible loss and you really won’t need it.

What’s with this Java -> Javascript stuff?

Why would somebody want to write a compiler for Java -> Javascript? One of the primary arguments for doing this is type-safety...which is fair enough I suppose - compile time checking is nice to have. You also get the benefit of native Java debugging tools, which despite the huge advances in client-side debugging in the past few years, Java’s debugging is still superior (mostly because of static typing). It’s really nice to be able to step through line by line, set breakpoints and inspect typed objects. However despite these benefits writing Java code that compiles into Javascript still feels weird.

The biggest reason for this awkwardness is that normally programmers write code in a language that is more expressive than your compile target, e.x. C/C++ compiles to assembly, Java to the JVM’s bytecode, CoffeeScript to Javascript etc. But with GWT your compile target is actually more expressive than the code you write and it’s not just a little bit more expressive, it’s a lot more expressive - which is a strange feeling indeed. Javascript has lambdas, prototypal inheritance, a less verbose syntax and a dynamic type system - all of which lead to a more expressive language, allowing you to do more with less code. Java on the other hand has none of these - expressing things that are normally trivial in Javascript (like a custom event system, currying etc.) can become a major chore (sometimes 4-5 classes or more) in GWT. Actually the whole process is rather akin to writing C/C++ code that compiles into Ruby or Python (ok, perhaps a slight exaggeration...but really only by a bit). But fear not, below are some tips to help make life easier for you in GWT-land.

Adjusting to life in GWT-land

  1. First use the gwt-mpv framework (http://www.gwtmpv.org/) - it will save your sanity and quite possibly your soul. One of our awesome developers, Stephen has created a very nice model, view, presenter framework on top of GWT that removes a huge amount of boilerplate. It has some really nice stuff like validation, two way data binding, code generation for tedious boilerplate and more. Your fingers and brain will thank you for saving them from the verbosity of vanilla GWT, trust me...I’m an engineer.

  2. Abandon the notion of separation of content (HTML), presentation (CSS), and behavior (Javascript) found in traditional front-end development. GWT, being a framework is opinionated and uses widgets instead. Widgets are generally responsible for all three of these things at once - they know their own CSS classes, keep their own data model and have event handlers. The idea behind this is you can just drop a widget onto any page and have it “just work” with no external dependencies. The disadvantage is the approach is inflexible if you want to change any of those three things while holding the others constant. In practice Widgets work quite well until you need to change or extend one that lives in somebody else’s jar - then they can quickly turn into a pain. So choose your imported widgets carefully or be prepared to fork things (or mash ctrl-c ctrl-v a lot).

  3. Even though you’re writing in Java - you’re still working on the client side. Custom events are great for decoupling and facilitating easy testing of your objects as you only have to talk to a single event handler instead of multiple objects. Don’t go overboard though, you still want to avoid having events that call events that call events etc. as you’ll end up jumping all over the place trying to find out what really was supposed to happen when the original event fired.

Wednesday, January 26, 2011

EMR/Hive: recovering a large number of partitions

If you try to run "alter table ... recover partitions" on a table with a large number of partitions, you may run into this error:


FAILED: Error in metadata: org.jets3t.service.S3ServiceException: Failed to sanitize XML document destined for handler class org.jets3t.service.impl.rest.XmlResponsesSaxParser$ListBucketHandler null 'null' -- ResponseCode: -1, ResponseStatus: null, RequestId: null, HostId: null
FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask


There's some discussion in the aws forums. The underlying cause is that it's running out of memory when trying to build the partition list.

A workaround is to increase the HADOOP_HEAPSIZE. This can be done by modifying hadoop-user-env.sh with an EMR bootstrap action. On an m1.large instance, 2G seems to do the trick for us.

Upload a script like the following somewhere in s3:



You can now run this bootstrap action as part of your job:

elastic-mapreduce --create --alive \
--name "large partitions..." --hive-interactive \
--num-instances 1 --instance-type m1.large \
--hadoop-version 0.20 \
--bootstrap-action s3://<bucket/path>/set-hadoop-heap.sh


You should now be able to load your partitions.

Wednesday, November 17, 2010

Spring NamespaceHandler debugging

While updating a library that used Spring yesterday, I began suffering the dreaded "Unable to locate Spring NamespaceHandler for XML schema namespace" exception.

The net is littered with threads about this exception, all of which end with "put the spring jars in your WEB-INF/lib directory".

I was fairly determined not to do this, as I frequently start a webapp with an embedded instance of Jetty, and use the Eclipse project's classpath for all of the dependencies. Since all of the jars are from the project's classpath, the webapp is always using exactly the same jars that Eclipse is for compiling your code, so you never have the two drift apart.

Too many times, after copying jars to WEB-INF/lib and forgetting about them, I'll upgrade a library, everything compiles fine, but spend an embarrassing amount of time wondering why it's not working in the webapp, before remembering the stale jar in the WEB-INF/lib directory.

Anyway, the real cause of the NamespaceHandler exception in my case was a buggy ClassLoader.getResources implementation.

The way spring works, when doing it's XML parsing/whatever magic, is when it sees "xmlns:tx=...", it wants a new NamespaceHandler that knows how to handle those tags.

To allow extensible NamespaceHandlers, Spring uses ClassLoader.getResources("META-INF/spring.handers") to get a list of all of the spring.handlers files across all of the jar files in the classloader. So if spring-core, spring-tx, spring-etc. all have spring.handlers with NamespaceHandlers in them, each file gets found and loaded.

Here's the rub: whatever Eclipse project classloader that had spring-tx on it only returned spring.handlers files from jars that already had classes loaded from them. The ClassLoader.getResources implementation would not look into jar files that was on its classpath, but had not yet been opened for loading classes from.

Of all things, adding:

Class.forName("org.springframework.transaction.support.TransactionTemplate");

Before spring was initialized fixed the NamespaceHandler error. Everything boots up correctly now.

While it took way too long to figure out, I'm pleased that I can continue using the Eclipse project classloader for the webapp I'm starting and avoid the annoying "copy jars to WEB-INF/lib" solution.

I'd like to know which classloader had the buggy getResources implementation, but I've already spent too much time on this so far.