Showing posts with label SmartFrog. Show all posts
Showing posts with label SmartFrog. Show all posts

Wednesday, 20 March 2013

Using Ansible for configuration management

In a previous life, I used to work as a research engineer in HP Labs Bristol (UK).

Over the 11 years or so I was there I worked on very cool projects (including a talking pot plant).

By far the project I had most fun with was FrameFactory, a Cloud Computing CGI rendering service that we rolled out as part of the SE3D showcase. FrameFactory was essentially a fairly complex distributed system with various moving parts that had to be configured and coordinated properly. To that end, we designed the system around SmartFrog and Anubis. This was handy as the engineering team for both of those projects was sitting in the nearby cubicles.

Anubis let us do 10 years ago what you now take from granted with Zookeeper, that is directory services and coordinating distributed system. SmartFrog is a very generic configuration management tool with a very nice DSL for expressing configuration data and a deployment engine for taking that and pushing onto remote nodes. I used SmartFrog in a lot of projects, and I even wrote a compiler plug-in for it that would auto-generate Java code from EMF. So as you can see it is very generic.

The downside of SmartFrog is that it is too generic, so if you wanted to use to it to install software and configure a Linux node with a particular role, you would still have to write a lot of low level drivers yourself. Which we did, and I remember writing SmartFrog scripts to deploy Xen vms and move them around using live migration. This was fun, but it felt that you had to write too much (Java) code. 

Forward to 2013, and I am finding myself working on another large scale distributed system that eventually has to be configured and deployed on Linux. I am not a Linux sysadmin, and I knew about Chef and Puppet, but given the amount of workload in that project (I have plenty of my plate), I got slightly worried just by reading the various ways Chef or Puppet (Solo, master, knife, etc..) could be used.

Then I came across Ansible and I was really pleased to find myself up and running within 30 minutes of reading the tutorial. The main web page does a good job of describing what Ansible does so I won't bother repeating it here. What I found is that Ansible makes it possible to turn Linux How-Tos documents (like how to enabled EPEL on CentOS) into workable, reusable scripts that are still very easy to understand.

Loading ....

Ansible has a comprehensive set of modules for installing packages, copying files over SSH, tweaking remote text files via its template mechanism. As a more complex example, I tried to use Ansible to see if I could use to deploy the lower stack required for high availability on Linux. Basically, this means installing Corosync, Pacemaker and ensuring that they are properly configured with the right IP addresses and so on.

What I had done so far was to capture all the steps required in documentation that could be reproduced (i.e. retyped), however I realized I could just come up with Ansible playbooks (recipes) to do the same and they would be just as clear. Here are a set of scripts which are sufficient to setup Corosync and Pacemaker working in UDP mode.

Loading ....

I also recently created playbooks to deploy CouchDB and Solr. I think tweaking those playbooks to setup replicated CouchDB or Solr clusters should not be too hard either.

Although I could (and should) give Puppet and Chef a closer look, I greatly value simplicity and I've been really impressed by Ansible's ease of use (the community is very active and friendly as well).

I've recommended that we use it as the Configuration Management solution for our current project.

Wednesday, 9 December 2009

Configuration management with Scala

Configuration management tools are essential things to have in your toolbox if your role is to manage large scale distributed IT systems. Very good free and open source solutions such as Puppet or SmarFrog exist for those familiar with Ruby or Java. It is not really my intention to implement yet another configuration management tool in Scala just for the sake of it. I know for a fact, having worked closely with the SmartFrog team in HP Labs Bristol, that they are rather complicated things to get right.

But that said....

The other day, I got thinking about Scala traits and the potential they offer for creating composable, refinable components that express configuration information and configuration logic. Other Scala features such as its type-safe nature and its package system could enable an elegant and simple way to represent configuration data and logic.

Let's start with a simple example that we will gradually extend. Let's assume I want to setup a cluster of nodes where each node is configured with an admin user account.

I start of with modelling a user with a User trait, a collection of properties such as the user's name, uid, etc.. The trait also contains (mock in this case) logic that can be executed to add or remove users on a computer. Notice how the user trait is itself composed from other traits.
trait Deployable{
  def deploy
  def undeploy
}

trait Named{
  var name : String = _
}

trait User extends Named with Deployable{
  var uid  : Short = _
  override def deploy = println ("useradd -u " + uid + " " + name)
  override def undeploy = println ("userdel -r " + name)
}

Users are to be deployed on nodes that I've modelled with a node trait (and once again that trait itself is composed from others). A node has an IP address and provides a convenience method to add things (such as a user) to it. Things added to the node will be deployed as the node is deployed by the hypothetical deployment runtime.
trait WithChildren{
  var includes : List[Any] = List()
  def contains (a : Any*) = a.foreach ( item => includes = item :: includes)
}

trait Node extends WithChildren with Deployable{
  var ip : String = _
  def deploy = println ("Actual logic to deploy children goes here.")
  def undeploy = println ("Actual logic to undeploy children goes here.")
}

trait Config extends WithChildren with Deployable{
  def deploy = println ("Actual logic to deploy nodes goes here.")
  def undeploy = println ("Actual logic to undeploy nodes goes here.")
}

The config trait models a collection of nodes. Again, in a real system, it would be the thing I actually pass to a deployment runtime for enaction.

With all the pieces in place, let's see a simple configuration:
object config1 extends Config{
  contains{
    new Node{
      ip = "192.168.1.1"
      contains{
        new User{
           name  = "admin"
           uid   = 102
        } 
      }
    }
  }
}

In the configuration above, I deploy the admin user on a single node. I can refine this a little bit by subclassing the node trait to create an AdminNode trait which includes the admin user by default.
object adminUser extends User{
  name ="admin"
  uid = 102
}

trait AdminNode extends Node{
 contains{
  adminUser
 }
}

In the snippet above, I've created an object adminUser which is a trait which has been instantiated. The adminUser can no longer be refined (subclassed) as Scala (unlike SmartFrog) is not a prototype based language. However the object can still be reused and composed into other traits. The other important thing to note is that, in Scala, when you create a trait, the logic that is executed when invoking its constructor is the entire body of the trait. So thanks to this feature, I can define new variables, change existing ones or invoke method calls within the curly braces without having to define an explicit constructor method as I would have to if using Java or Groovy.

object config2 extends Config{
  contains{
    new AdminNode{ ip = "192.168.1.2" }
  } 
}


If you want to modify a particular instance of AdminNode in place let's say to add another user account, it is also easily done.

object config3 extends Config{
    contains{
      new AdminNode{
       ip = "192.168.1.3"
       contains{
        new User{
         name  = "demo"
         uid = 102
        }
       }
      }
   }
}  


One of the benefits of using Scala directly to write the configuration is that I can use constructs such as loops to create or modify the data. For instance, let's create a set of admin nodes from a list of IP addresses.
object config4 extends Config{
  List("192.168.1.2","192.168.1.3","192.168.1.4").foreach{ addr=>
    contains{
     new AdminNode{ ip = addr}
    }
  }
}


Just as I modelled users, I can also model applications running on nodes (again by composing traits). In the example below I model generic applications installed through packages (via a package manager a la apt-get) and controlled via Linux services.

trait Package extends Deployable{
  var packages : List[String] = List()
  override def deploy : Unit = println (packages.foreach(s => "apt-get install " + s))
  override def undeploy : Unit = println (packages.foreach(s => "apt-get remove " + s))
}

trait Services extends Runnable{
  var services : List[String] = List()
  override def start : Unit = println (services.foreach(s => "/etc/init.d/" + s + " start"))
  override def stop : Unit = println (services.foreach(s => "/etc/init.d/" + s + " stop"))
}


Using those traits, I can then model an application such as an Apache Web Server, or refine an Apache Web server into a Django application server running as an apache module.
trait WebServer{
 var port = 8080
}

trait Apache2 extends Services with Package with WebServer{
 packages += "apache2"
 services += "apache2"
 port = 80
}

trait Django extends Apache2{
 packages = "libapache2-mod-python" :: "python-django" :: packages
}

I can then include instances of Apache2 or Django in my hypothetical cluster.
object config5 extends Config{
    contains(
     new AdminNode{
      ip = "192.168.1.3"
      contains{
             new Apache2{ port = 8182 }
             }
     },    
     new AdminNode{
      ip = "192.168.1.4"
      contains{
             new Django{ port = 87 }
             }
     }
    )
}

As the configuration information is written directly in Scala, I automatically gain access to interesting features:
  • the compiler highlights syntax errors in the description
  • I can use IDE for syntax highlighting, auto completion and re-factoring
  • descriptions can be organised into packages and imported as required.

Obviously, as a thought experiment, this ought to be taken with a pinch of salt. I've only touched on some of the language features that are appropriate for expressing composable and reusable models of configuration data. I have not tried (and most likely wont try) to implement a distributed deployment engine that could deploy such configuration descriptions.