2011-07-06

Neo4jPHP beta released

I've been working on a PHP client for the Neo4j REST API for little while. I think it's ready for some real-life testing. The beta version is available here: https://github.com/jadell/Neo4jPHP/tarball/0.0.1-beta

Features:
  • Developed against the Neo4j 1.4 milestone releases
  • Simple, object-oriented API
  • Almost complete REST API coverage
  • Indexing of nodes and relationships, including exact match and query support
  • Cypher queries (thanks to Jacob Hansson)
  • Traversal support, including paged traversals
  • Lazy-loading of node and relationship data

Hopefully coming soon:
  • Client-side caching
  • Batch operations

There are some usage examples included.

It's a beta release, so please be gentle (on me, that is; be as rough as you want with the code.) If anyone finds any bugs or has feature requests, please use the GitHub issues page.


Update 2011-07-08:I would be remiss if I didn't mention the other PHP Neo4j REST client available at https://github.com/tchaffee/Neo4J-REST-PHP-API-client. It is built for the current stable release of Neo4j 1.3.

2011-06-27

Path finding with Neo4j

In my previous post I talked about graphing databases (Neo4j in particular) and how they can be applied to certain classes of problems where data may have multiple degrees of separation in their relationships.

The thing that makes graphing databases useful is the ability to find relationship paths from one node to another. There are many algorithms for finding paths efficiently, depending on the use case.

Consider a graph that represents a map of a town. Every node in the graph represents an intersection of two or more streets. Each direction of a street is represented as a different relationship: two-way streets will have two relationships between the same two nodes, one pointing from the first node to the second, and the other pointing from the second node to the first. One-way streets will have a single relationship. Each street relationship will also have a distance property. Here is what one such map might look like:
Every street is two-way, except for Market Alley and Park Drive. The numbers in parentheses represent the street distance in hundreds of yards (i. e. 3 = 300 yards.)

Here is the code to create this graph in the database, using the REST client I've been working on (connection initialization is not shown):
$intersections = array(
    "First & Main",
    "Second & Main",
    "Third & Main",
    "First & Oak",
    "Second & Oak",
    "Third & Market",
);

$streets = array(
    // start, end, [direction, distance, name]
    array(0, 1, array('direction'=>'east', 'distance'=>2, 'name'=>'Main St.')),
    array(0, 3, array('direction'=>'south', 'distance'=>2, 'name'=>'First Ave.')),
    array(0, 4, array('direction'=>'south', 'distance'=>3, 'name'=>'Park Dr.')),

    array(1, 0, array('direction'=>'west', 'distance'=>2, 'name'=>'Main St.')),
    array(1, 2, array('direction'=>'east', 'distance'=>2, 'name'=>'Main St.')),
    array(1, 4, array('direction'=>'south', 'distance'=>2, 'name'=>'Second Ave.')),

    array(2, 1, array('direction'=>'west', 'distance'=>2, 'name'=>'Main St.')),
    array(2, 5, array('direction'=>'south', 'distance'=>1, 'name'=>'Third Ave.')),

    array(3, 0, array('direction'=>'north', 'distance'=>2, 'name'=>'First Ave.')),
    array(3, 4, array('direction'=>'east', 'distance'=>2, 'name'=>'Oak St')),

    array(4, 3, array('direction'=>'west', 'distance'=>2, 'name'=>'Oak St.')),
    array(4, 1, array('direction'=>'north', 'distance'=>2, 'name'=>'Second Ave.')),

    array(5, 2, array('direction'=>'north', 'distance'=>1, 'name'=>'Third Ave.')),
    array(5, 4, array('direction'=>'west', 'distance'=>2, 'name'=>'Market Alley')),
);

$nodes = array();
$intersectionIndex = new Index($client, Index::TypeNode, 'intersections');
foreach ($intersections as $intersection) {
    $node = new Node($client);
    $node->setProperty('name', $intersection)->save();
    $intersectionIndex->add($node, 'name', $node->getProperty('name'));
    $nodes[] = $node;
}

foreach ($streets as $street) {
    $start = $nodes[$street[0]];
    $end = $nodes[$street[1]];
    $properties = $street[2];
    $start->relateTo($end, 'CONNECTS')
        ->setProperties($properties)
        ->save();
}
First, we set up a list of the intersections on the map, and also a list of the streets that connect those intersections. Each street will be a relationship that has data about the streets direction, distance and name.

Then, we create each intersection node. Each node is also added to an index. Indexes allow stored nodes to be found quickly. Nodes and relationships can be indexed, and indexing can occur on any node or relationship field. You can even index on fields that don't exist on the node or relationship. We'll use the index later to find the start and end points of our path by name.

Finally, we create a relationship for every street. Relationships can have arbitrary data attached to them, just like nodes.

Now we create way to find and display driving directions from any intersection to any other intersection:
$turns = array(
    'east' => array('north' => 'left','south' => 'right'),
    'west' => array('north' => 'right','south' => 'left'),
    'north' => array('east' => 'right','west' => 'left'),
    'south' => array('east' => 'left','west' => 'right'),
);

$fromNode = $intersectionIndex->findOne("Second & Oak");
$toNode = $intersectionIndex->findOne("Third & Main");

$paths = $fromNode->findPathsTo($toNode, 'CONNECTS', Relationship::DirectionOut)
    ->setMaxDepth(5)
    ->getPaths();

foreach ($paths as $i => $path) {
    $path->setContext(Path::ContextRelationship);
    $prevDirection = null;
    $totalDistance = 0;

    echo "Path " . ($i+1) .":\n";
    foreach ($path as $j => $rel) {
        $direction = $rel->getProperty('direction');
        $distance = $rel->getProperty('distance') * 100;
        $name = $rel->getProperty('name');

        if (!$prevDirection) {
            $action = 'Head';
        } else if ($prevDirection == $direction) {
            $action = 'Continue';
        } else {
            $turn = $turns[$prevDirection][$direction];
            $action = "Turn $turn, and continue";
        }
        $prevDirection = $direction;
        $step = $j+1;
        $totalDistance += $distance;

        echo "\t{$step}: {$action} {$direction} on {$name} for {$distance} yards.\n";
    }
    echo "\tTravel distance: {$totalDistance} yards\n\n";
}
Note that in the `findPathsTo` call, we specify that we want only outgoing relationships. This is how we prevent getting directions that may send us the wrong way down a one-way street. We also set a max depth, since the default search depth is 1. Also, note the `setContext` call which tells the path that we want the `foreach` to loop through the relationships, not the nodes. Running this produces the following directions:
Path 1:
    1: Head north on Second Ave. for 200 yards.
    2: Turn left, and continue west on Main St. for 200 yards.
    Travel distance: 400 yards
We only got back one path because, by default, the path finding algorithm finds the shortest path, meaning the path with the least number of relationships. If we switch the to and from intersections, we can get directions back to our starting point:

Path 1:
    1: Head west on Main St. for 200 yards.
    2: Turn left, and continue south on Second Ave. for 200 yards.
    Travel distance: 400 yards

Path 2:
    1: Head south on Third Ave. for 100 yards.
    2: Turn right, and continue west on Market Alley for 200 yards.
    Travel distance: 300 yards
We got back two paths this time, but one of them is longer than the other. How can this be if we are using the shortest path algorithm? The answer is that shortest path refers only to the number of relationships in the path. We didn't tell our path finder to pay attention to the cost of one path over another. Fortunately, we can do that quite easily by telling the `findPathsTo` call to use a property of each relationship when determining which path is the least expensive:
$paths = $fromNode->findPathsTo($toNode, 'CONNECTS', Relationship::DirectionOut)
    ->setAlgorithm(PathFinder::AlgoDijkstra)
    ->setCostProperty('distance')
    ->setMaxDepth(5)
    ->getPaths();
We're telling the path finder to use Dijkstra's algorithm, an algorithm to find the lowest cost path, which is different from the shortest path. Since our map uses distance, the cost of a path is the total of all the distance properties of the relationships in that path. Running the script now gives us only one path, the path with the least distance:
Path 1:
    1: Head south on Third Ave. for 100 yards.
    2: Turn right, and continue west on Market Alley for 200 yards.
    Travel distance: 300 yards
If there were another path with the same total distance, it would also be displayed. Path cost ignores the number of relationships in the path. If there were a path with 3 relationships totaling 300 yards, and another with 1 relationship totaling 300 yards, both of them would be displayed.

Neo4j has several path finding algorithms built in. Some of them are:
  • shortest path - find paths with the fewest relationships
  • dijkstra - find paths with the lowest cost
  • simple path - find paths with no repeated nodes
  • all - find all paths between two nodes
Neo4j also comes with an expressive traversal API that allows you to create your own path finding algorithms for domain specific needs.

2011-06-15

Neo4j for PHP

Update 2011-09-14: I've put a lot of effort into Neo4jPHP since this post was written. It's pretty full-featured and covers almost 100% of the Neo4j REST interface. Anyone interested in playing with Neo4j from PHP should definitely check it out. I would love some feedback!

Lately, I've been playing around with the graph database Neo4j and its application to certain classes of problems. Graph databases are meant to solve problems in domains where data relationships can be multiple levels deep. For example, in a relational database, it's very easy to answer the question "Give me a list of all actors who have been in a movie with Kevin Bacon":
> desc roles;
+-------------+--------------+------+-----+---------+----------------+
| Field       | Type         | Null | Key | Default | Extra          |
+-------------+--------------+------+-----+---------+----------------+
| id          | int(11)      | NO   | PRI | NULL    | auto_increment |
| actor_name  | varchar(100) | YES  |     | NULL    |                |
| role_name   | varchar(100) | YES  |     | NULL    |                |
| movie_title | varchar(100) | YES  |     | NULL    |                |
+-------------+--------------+------+-----+---------+----------------+

> SELECT actor_name, role_name FROM roles WHERE movie_title IN (SELECT DISTINCT movie_title FROM roles WHERE actor_name='Kevin Bacon')
Excuse my use of sub-queries (re-write it to a JOIN in your head if you wish.)

But suppose you want to get the names of all the actors who have been in a movie with someone who has been in a movie with Kevin Bacon. Suddenly, you have yet another JOIN against the same table. Now add a third degree: someone who has been in a movie with someone who has been in a movie with someone who has been in a movie with Kevin Bacon. As you continue to add degrees, the query becomes increasingly unwieldy, harder to maintain, and less performant.

This is precisely the type of problem graph databases are meant to solve: finding paths between pieces of data that may be one or more relationships removed from each other. They solve it very elegantly by modeling domain objects as graph nodes and edges (relationships in graph db parlance) and then traversing the graph using well-known and efficient algorithms.

The above example can be very easily modeled this way: every actor is a node, every movie is a node, and every role is a relationship going from the actor to the movie they were in:
Now it becomes very easy to find a path from a given actor to Kevin Bacon.

Neo4j is an open source graph database, with both community and enterprise licensing structures. It supports transactions and can handle billions of nodes and relationships in a single instance. It was originally built to be embedded in Java applications, and most of the documentation and examples are evidence of that. Unfortunately, there is no native PHP wrapper for talking to Neo4j.

Luckily, Neo4j also has a built-in REST server and PHP is very good at consuming REST services. There's already a good Neo4j REST PHP library out there, but I decided to write my own to get a better understanding of how the REST interface actually works. You can grab it here and all the code examples below are written using it. The concepts can easily be ported to any Neo4j REST client.

First, we need to initialize a connection to the database. Since this is a REST interface, there is no persistent connection, and in fact, no data communication happens until the first time we need to read or write data:
use Everyman\Neo4j\Client,
    Everyman\Neo4j\Transport,
    Everyman\Neo4j\Node,
    Everyman\Neo4j\Relationship;

$client = new Client(new Transport('localhost', 7474));
Now we need to create a node for each of our actors and movies. This is analogous to the INSERT statements used to enter data into a traditional RDBMS:
$keanu = new Node($client);
$keanu->setProperty('name', 'Keanu Reeves')->save();
$laurence = new Node($client);
$laurence->setProperty('name', 'Laurence Fishburne')->save();
$jennifer = new Node($client);
$jennifer->setProperty('name', 'Jennifer Connelly')->save();
$kevin = new Node($client);
$kevin->setProperty('name', 'Kevin Bacon')->save();

$matrix = new Node($client);
$matrix->setProperty('title', 'The Matrix')->save();
$higherLearning = new Node($client);
$higherLearning->setProperty('title', 'Higher Learning')->save();
$mysticRiver = new Node($client);
$mysticRiver->setProperty('title', 'Mystic River')->save();
Each node has `setProperty` and `getProperty` methods that allow storing arbitrary data on the node. No server communication happens until the `save()` call, which must be called for each node.

Linking an actor to a movie means setting up a relationship between them. In RDBMS terms, the relationship takes the place of a join table or a foreign key column. In the example, the relationship always starts with the actor pointing to the movie, and is tagged as an "acted in" type relationship:
$keanu->relateTo($matrix, 'IN')->save();
$laurence->relateTo($matrix, 'IN')->save();

$laurence->relateTo($higherLearning, 'IN')->save();
$jennifer->relateTo($higherLearning, 'IN')->save();

$laurence->relateTo($mysticRiver, 'IN')->save();
$kevin->relateTo($mysticRiver, 'IN')->save();
The `relateTo` call returns a Relationship object, which is like a node in that it can have arbitrary properties stored on it. Each relationship is also saved to the database.

The direction of the relationship is totally arbitrary; paths can be found regardless of which direction a relationship points. You can use whichever semantics make sense for your problem domain. In the example above, it makes sense that an actor is "in" a movie, but it could just as easily be phrased that a movie "has" an actor. The same two nodes can have multiple relationships to each other, with different directions, types and properties.

The relationships are all set up, and now we are ready to find links between any actor in our system and Kevin Bacon. Note that the maximum length of the path is going to be 12 (6 degrees of separation multiplied by 2 nodes for each degree; an actor node and a movie node.)
$path = $keanu->findPathsTo($kevin)
    ->setMaxDepth(12)
    ->getSinglePath();

foreach ($path as $i => $node) {
    if ($i % 2 == 0) {
        echo $node->getProperty('name');
        if ($i+1 != count($path)) {
            echo " was in\n";
        }
    } else {
        echo "\t" . $node->getProperty('title') . " with\n";
    }
}
A path is an ordered array of nodes. The nodes alternate between being actor and movie nodes.

You can also do the typical "find all the related movies" type queries:
echo $laurence->getProperty('name') . " was in:\n";
$relationships = $laurence->getRelationships('IN');
foreach ($relationships as $relationship) {
    $movie = $relationship->getEndNode();
    echo "\t" . $movie->getProperty('title') . "\n";
}
`getRelationships` returns an array of all relationships that a node has, optionally limiting to only relationships of a given type. It is also possible to specify only incoming or outgoing relationships. We've set up our data so that all 'IN' relationships are from an actor to a movie, so we know that the end node of any 'IN' relationship is a movie.

There is more available in the REST interface, including node and relationship indexing, querying, and traversal (which allows more complicated path finding behaviors.) Transaction/batch operation support over REST is marked "experimental" for now. I'm hoping to add wrappers for more of this functionality soon. I'll also be posting more on the different types of problems that can be solved very elegantly with graphing databases.

The next post dives into a bit more detail about path-finding and some of the different algorithms available.

2011-05-16

Logging User Sessions Across Requests

When tracking down a bug, few things aid the process more than good application logs. And this is especially true when the bug has already escaped to your production systems. In these cases, you may not be able to inject test data into your database, cowboy-code in some `var_dump`s or perform the same action over and over again. Logs can save a lot of time and effort when tracking down issues in a live application.

But logs also have a few shortcomings. In an active application with multiple users, all the log messages for every user are piled together in the same log file. If the application is multi-user (like most web applications), log messages from different requests can be interwoven. This makes it a nightmare to try and track the log messages for a single user request, let alone messages across several requests by the same user.

One way to handle this is to put a request-specific identifier in every log message. But I shouldn't have to remember to append or prepend the identifier to their log messages. I'd rather have it happen automatically, without me or my teammates having to think about it.

Here's a method we've been using to try and untangle the mess and retain the usefulness of our logs. The code uses Zend's logging component, but can easily be adapted to other log systems.

First, we start up a log and tell it how to format each log line.
$format = "%timestamp% [%logId%]: %message%" . PHP_EOL;
$formatter = new Zend_Log_Formatter_Simple($format);

$writer = new Zend_Log_Writer_Stream('/path/to/application.log');
$writer->setFormatter($formatter);

$log = new Zend_Log();
$log->addWriter($writer);
In the format string, names wrapped in "%" signs are log variables that will be filled in when a message is logged. "timestamp" and "message" are provided by Zend and will be filled with the appropriate values when the logging call is made. "logId" is a custom symbol, which means we need to tell the log object what value to give it.

In order to track all the log messages across a single user request, we need to pick a unique value for the log id:
// The last 8 characters of a uniqid are unique enough for our purposes
$logId = substr(uniqid(), -8);
$log->setEventItem('logId', $logId);
Now we wait for a few users to make requests at the same time:
// A web request
$log->info('The user made some request');
$log->info('Here is some output of that request: '.$output);

// A different, simultaneous request from another user
$log->info('The user made some request');
$log->info('Here is some output of that request: '.$output);

// Another request from the first user
$log->info('Another user request');
Here's the log output:
2011-05-16 19:04 [a73fe12a]: The user made some request
2011-05-16 19:04 [24f90ae1]: The user made some request
2011-05-16 19:04 [24f90ae1]: Here is some output of that request: Hello, World
2011-05-16 19:04 [a73fe12a]: Here is some output of that request: Hello, World
2011-05-16 19:04 [feb40a52]: Another user request
Now we can easily track a single request's messages in the log, even if multiple requests happen at the same time.

But we can do better. Often times, bugs are not the result of a single user request, but of a sequence of requests. If we only track logs on a per request basis, we haven't helped ourselves track request-spanning bugs. Luckily, we already have a value that uniquely identifies a user across multiple requests: their PHP session id. We can use it, in addition to a request identifier:
// The first 8 characters of the session id are unique enough
session_start();
$sessId = substr(session_id(), 0, 8);
$requestId = substr(uniqid(), -8);
$logId = $sessId.'-'.$requestId;
$log->setEventItem('logId', $logId);
And here's what our logs look like now:
2011-05-16 19:04 [b74be12f-a73fe12a]: The user made some request
2011-05-16 19:04 [075aef5e-24f90ae1]: The user made some request
2011-05-16 19:04 [075aef5e-24f90ae1]: Here is some output of that request: Hello, World
2011-05-16 19:04 [b74be12f-a73fe12a]: Here is some output of that request: Hello, World
2011-05-16 19:04 [b74be12f-feb40a52]: Another user request
And there it is. We can use the log id in useful ways, such as displaying it to the user with a "contact tech support" message. When a user calls or emails and says their "support id" is "b74be12f-a73fe12a" we not only have the log messages for the request that gave them the error, but all the messages for their entire session:
# All the request log messages
> grep 'b74be12f-a73fe12a' /path/to/application.log
# All the session log messages
> grep 'b74be12f-' /path/to/application.log

Here's all the code together:
$format = "%timestamp% [%logId%]: %message%" . PHP_EOL;
$formatter = new Zend_Log_Formatter_Simple($format);

$writer = new Zend_Log_Writer_Stream('/path/to/application.log');
$writer->setFormatter($formatter);

$log = new Zend_Log();
$log->addWriter($writer);

session_start();
$sessId = substr(session_id(), 0, 8);
$requestId = substr(uniqid(), -8);
$logId = $sessId.'-'.$requestId;
$log->setEventItem('logId', $logId);

$log->info('The user made some request');
$log->info('Here is some output of that request: '.$output);

2011-05-03

Interview with Bulat Shakirzyanov

Today on Cal Evans' "Voices of the ElePHPant" there was an interview with Bulat Shakirzyanov. His comments about mocking PHP's built-in functions reminded me of this PHP built-in mocking library I had made a little while ago, then forgotten about. Bulat's example talks about `mkdir`, which the built-in mock library would handle, but for filesystem specific functionality, I also highly recommend vfStream. It allows you to create a virtual filesystem, comeplete with directory hierarchy and files with contents, all from within a unit test, and all without needing an actual filesystem.

2011-04-22

Tropo-Turing Collision: Voice Chat Bots with Tropo

On the first evening of PHPComCon this year, Tropo sponsored a beer-and-pizza-fueled hackathon. Having never heard of Tropo before, I decided to check it out.

Tropo provides a platform for bridging voice, SMS, IM and Twitter data to create interactive applications in Javascript, PHP, Python and a couple other languages. An example would be the call-center navigation menus with which many people are familiar, or a voicemail system that sends a text message when a new message is received. But those are only the tip of the iceberg.

To explore a little deeper, I thought it would be neat to create a voice chat-bot. The idea would be that a caller could talk to an automated voice in a natural way, and the voice would respond in a relevant way and push the conversation along. Since I wasn't aiming for Turing test worthiness, a good starting point was ELIZA, one of the first automated chat-bots. I thought it was a pretty ambitious project, but it turns out that Tropo's system handles all of the functionality right out of the gate.

The docs do a great job of explaining setting up an account and creating an application, so I'm going to jump right into the code (in PHP).

I like to keep my code relatively clean and organized, with functions and classes in their own files. So the first thing I did was create a hosted file called "Eliza.php" with the following contents:
class Eliza
{
  public function respondsTo($statement="")
  {
     $responses = array(
       "0" => "one",
       "1" => "two",
       "2" => "three",
       "3" => "four",
       "4" => "five",
       "5" => "six",
       "6" => "seven",
       "7" => "eight",
       "8" => "nine",
       "9" => "zero",
     );

     $response = "Please pick a number from 0 to 9";
     if (isset($responses[$statement])) {
       $response = $responses[$statement];
     }

     return $response;
  }

  public function hears($prompt)
  {
    $result = ask($prompt, array(
      "choices" => "[1 DIGIT]"
    ));
    $statement = strtolower($result->value);
    return $statement;
  }
}
It's important to note that the opening <?php and closing ?> should be left out of this file, or the next bit will fail.

The main functionality of the application is in the `Eliza` class. Eliza will translate a user input string into a response, which it will then use to prompt the user. The `ask()` function in the `hears()` method is functionality provided by Tropo's system that takes care of the text-to-speech and speech-to-text aspect of prompting the caller, and then waits for the caller to respond. The `choices => [1 DIGIT]` option to `ask()` hints that we expect the user to respond to our prompt with a single 0-9 character.

Next, I created a file called "chatbot.php" with the following contents:
<?php
$url = "http://hosting.tropo.com/00001/www/Eliza.php";
$ElizaFile = file_get_contents($url);
eval($ElizaFile);

$eliza = new Eliza();
$statement = "";
do {
  $prompt = $eliza->respondsTo($statement);
  $statement = $eliza->hears($prompt);
  _log("They said ".$statement);
} while (true);
?>
This is the entry script for the application, which can be set on the "Application Settings" page. Unfortunately, it does not look like Tropo supports `require` and `include` in their system. Fortunately, what they do provide are URLs to download the contents of any hosted file via simple `file_get_contents`. So we "inlcude" our Eliza.php file by downloading its contents, then `eval`ing them into the running scripts scope. Note: Yes I could have just written the contents of Eliza.php into the chatbot.php file, but a) it was more fun to try and find a way around that limitation :-) and b) many developers separate their code this way to keep it clean, encapsulated and reusable and this demonstrates a way to accomplish that.

The code is fairly self-explanatory: "include" and instantiate Eliza, then enter a prompt-respond loop which will last until the caller hangs up. `_log()` outputs to Tropo's built in application debugger.

I have to say congratulations to Tropo for creating a platform that makes all this easy. I had this up and running (except the file include portion) in about 15 minutes. You can try this out for yourself by calling (919) 500-7747.

So now that prompt-response was working, I could get started on making a real Eliza chat-bot that would parse the caller's speech and provide an Eliza-like response:

Caller: I'm building a voice chat application.
Eliza: Tell me more about voice chat application.
Caller: You're it!
Eliza: How do you feel about that?
Caller: I think it's pretty neat.

There are probably a hundred Eliza implementations on the web, and at least a dozen are written in PHP. I grabbed the first one I found, shoved it into the `Eliza::respondsTo()` method, and called up my application.

And this is where I hit the iceberg. In order to accomplish what I wanted, I needed the `ask()` function to be able to accept and parse any spoken words into a string. This meant getting rid of the `choices => [1 DIGIT]` line. As soon as I called, Eliza started prompting me over and over for input, until Tropo's system killed the loop and hung up.

Luckily, a Tropo guy (who was great to talk to, but who's name I have unfortunately forgotten) informed me that `ask()` works by using the `choices` option as training for the speech-to-text parser. There is a way to do generalized speech-to-text but it is incredibly processor intensive, and would have to make use of an asynchronous call. Later that evening, @akalsey confirmed that the current state of the technology (everyone's, not just Tropo's) makes generalized real-time speech-to-text processing impossible. So the dream of speaking with Eliza instead of just IMing with her dies unfulfilled, or is at least put on the shelf until technology catches up with ideas.

I did learn some about Tropo's service, though, so I count the hackathon successful. Thanks again, Tropo, for the great new tool!

2011-04-14

Syntactic sugar for Vows.js

We end up writing quite a few full-stack integration tests. As the workflows for our projects get longer and more complex, the corresponding Zombie and Vows tests also get longer and more complex. Soon enough, a single test can be nested 15 or 20 levels deep, with a mix of anonymous functions and commonly-used macros.

Needless to say, these tests can become monsters to try and read through and understand:
var vows = require('vows');

vows.describe("Floozle Workflow").addBatch({
    "Go to login page" : {
        topic : function () {
            zombie.visit("http://widgetfactory.com", this.callback);
        },
        "then login" : {
            topic : function (browser) {
                browser
                  .fill("#username", "testuser")
                  .fill("#password", "testpass")
                  .pressButton("Login", this.callback);
            },
            "then navigate to Floozle listing" : {
                topic : function (browser) {
                    browser.click("Floozles", this.callback);
                },
                "should be on the Floozle listing" : function (browser) {
                    assert.include(browser.text("h3"), "Floozles");
                },
                "should have a 'Create Floozle' link" : function (browser) {
                    assert.include(browser.text("Create Floozle"));
                },
                "then click 'Create Floozle'" : {
                    topic : function (browser) {
                        browser.click("Create Floozle");
                    },
                    "then fill out Floozle form" : {
                        topic : function (browser) {
                            browser
                              .fill("#name", "Klaxometer")
                              .fill("#whizzbangs", "27")
                              .choose("#rate", "klaxes per cubic freep")
                              .pressButton("Save", this.callback);
                        },
                        // .... Continue on from there, with more steps
                    }
                }
            }
        }
    }
}).export(module);
So we created prenup, a syntactic sugar library for easily creating Vows tests. Prenup provides a fluent interface for generating easy to read test structures as well as the ability to reuse and chain testing contexts.

Here is the above test written with prenup:
var vows = require('vows'),
    prenup = require('prenup');

var floozleWorkflow = prenup.createContext(function () {
    zombie.visit("http://widgetfactory.com", this.callback);
})
.sub("then login", function (browser) {
    browser
      .fill("#username", "testuser")
      .fill("#password", "testpass")
      .pressButton("Login", this.callback);
})
.sub("then navigate to Floozle listing", function (browser) {
    browser.click("Floozles", this.callback);
})
.vow("should be on the Floozle listing", function (browser) {
    assert.include(browser.text("h3"), "Floozles");
})
.vow("should have a 'Create Floozle' link", function (browser) {
    assert.include(browser.text("Create Floozle"));
})
.sub("then click 'Create Floozle'", function (browser) {
    browser.click("Create Floozle");
})
.sub("then fill out Floozle form", function (browser) {
    browser
      .fill("#name", "Klaxometer")
      .fill("#whizzbangs", "27")
      .choose("#rate", "klaxes per cubic freep")
      .pressButton("Save", this.callback);
})
.root();

vows.describe("Floozle Workflow").addBatch({
    "Go to login page" : floozleWorkflow.seal()
}).export(module);
Every `sub()` call generates a new sub-context, and every `vows()` call attaches a vow assertion to the most recent context. `parent()` can be called to pop back up the chain and attach parallel contexts. If a context is saved, it can be reused in other chains. When `seal()` is called, it will generate the same structure as the original test.

More examples, including branching and reusing contexts, are given in the documentation. Prenup is available on NPM via `npm install prenup`