Showing posts with label clojure. Show all posts
Showing posts with label clojure. Show all posts

Wednesday, September 12, 2018

Manufacturing Creativity

Previously, I've attempted to convince you that making software is a creative act, and I explored the implications for pursuing and managing software engineering. (By the way, science and engineering are also creative acts, and a great exploration of that idea is "Better Science Through Art" by Richard P. Gabriel and Kevin J. Sullivan. Love that paper.)

I've been thinking a lot lately about creativity and how it can be encouraged (even manufactured?). I've also been thinking quite a bit about why people do (or do not) take on ambitious projects, and how to survive a years long ambitious project. I've learned some very interesting things that some day I may write about, but I'd like to share what I've learned about creativity.

What I've discovered about being creative is that even from people in very different lines of work (actors, writers, artists, programmers, scientists, investors) there's a surprising amount of agreement about how it works. I've also discovered that it is not an innate talent that some people have and some do not. Everyone has the tools to be creative.

In many ways this goes all the way back to the very first Clojure Conj in October of 2010. Rich Hickey gave a talk titled "Step Away from the Computer"...actually, it had three titles, and it is best known by one of its other titles "Hammock-Driven Development." I was there in person. I came away with the mistaken impression that the talk was about writing software and solving technical problems. I now know that making software is a creative act, and Rich's talk was about how to be creative.

For Rich, the engine of creativity is the "background mind," which is in contrast to the "waking mind." Your waking mind is your normal mode of operation. It is good at analyzing and thinking critically, but can be too tactical and get stuck in local maxima. Your background mind is good at making connections, thinking abstractly, and synthesizing. It can make the leap past local maxima, unfortunately your background mind cannot be tasked directly. However, you can task it indirectly by obsessively thinking and reading about a particular problem, and, though you can activate it other ways, it is easiest to activate it by sleeping, or relaxing and simulating sleeping (i.e. using a hammock).

So, creativity is an indirect process of a relaxed mental mode that you task by obsessively thinking about a problem, and whose products you only filter after the fact with your normal critical-analytical mental mode. Now here's the surprising part, almost everyone who attempts to describe their creative process describes it similarly. In his essay, "The Top Ideas in Your Mind," Paul Graham says:

Everyone who's worked on difficult problems is probably familiar with the phenomenon of working hard to figure something out, failing, and then suddenly seeing the answer a bit later while doing something else. There's a kind of thinking you do without trying to. I'm increasingly convinced this type of thinking is not merely helpful in solving hard problems, but necessary. The tricky part is, you can only control it indirectly.

John Cleese gave a talk on creativity, and he called the background mind "open mode" and the waking mind "closed mode." In your open mode, you are relaxed, less purposeful, curious, and a bit playful. In your closed mode, you are active, determined, and have a critical eye.

George Land is a business man who investigated how to stimulate and direct creativity. He found there are two kinds of thinking: divergent and convergent. Divergent thinking is creating new ideas. Convergent thinking is judging and evaluating ideas. He did a longitudinal study that found that 98% of 5 year olds exhibit divergent thinking, 30% of 10 year olds, 12% of 15 year olds, and only 2% of adults think divergently. As a person gets older, he or she is taught to use both divergent and convergent thinking at the same time. The result is one criticizes and judges ideas before they can fully develop.

Of peculiar interest to me has been what independent game designer Jonathan Blow—who worked on his successful and influential game Braid for 3.5 years—has said about creativity and surviving ambitious projects. (Maybe someday Rich will talk about how he survived his own ambitious projects: how to maintain motivation day-to-day, how to fund it, how to plan and pace it, how to finish it.) The thoughts about ambitious projects are for another time, but what he says about creativity should be familiar by now. Blow says metaphysically you may not buy into the Greek concept of the Muse—nor may he—but functionally it is real. Creativity feels like something external, and you have to get yourself into a relaxed mode to provide opportunity for new ideas, though you cannot guarantee anything.

Tools and Techniques


I hope to find more resources on direct techniques for stimulating creative (e.g. instead of thinking about solving a problem think about how to make in worse and avoid that), but for now I've found a lot of agreement about how to encourage creativity in an indirect way.

Obsess about your problem. If your subconscious mind (or unconscious mind or background mind or whatever you want to call it) is going to solve a problem for you, then it needs information. Rich has a lot of great advice about this. Write down your problem. Write down what you know. Write down what you don't know. Read about your problem. Read about related problems. Pick apart other solutions. Paul Graham in "The Top Idea in Your Mind" says, "It's hard to do a really good job on anything you don't think about in the shower."

Relax. For Rich this is lying in a hammock and focusing, thinking through all the information you've loaded into your mind. For Blow, a relaxed state of mind is really a pretty active body. He likes to find something purely physical that he can enjoy, like going to a club and dancing. Cleese creates an oasis blocking off time and setting aside other concerns. He gives himself enough time that he can work through all the TODOs that pop into his head. He writes them down for later, and gets back to being relaxed and playful.

Pace yourself. Cleese recommends, if you're going to try to set aside time for creativity, to limit it to no more than an hour and a half, because you'll need a break. If you need more time, then do it again the next day.

Be Playful. George Land found that children are more creative. Cleese finds being in a playful mood conducive to creativity, especially when collaborating with others. Play, imagination, daydreaming all come from or lead to a relaxed state of mind, which accesses your creative mechanism.

Write things down. Rich is big on this. There are several benefits: it helps you think thoroughly, it helps you remember things, it is easy to skim for recall.

Gently keep your mind focused. Cleese says to be successful you must keep your mind gently around the problem. You may wander off, but gently come back to it. Rich uses hammock time not just to relax, but to recall information. Touch each fact with your mind to keep it fresh, and to make it interesting to your background mind.

Have a dogged persistence. Cleese sticks with an problem, and doesn't just take the first idea he comes up with. Sometimes a creative breakthrough requires persisting through the discomfort, even slight anxiety, of an unsolved problem. Rich reminds us that since this is an indirect process it may take days, months, or years for a solution to come.

Anti-techniques


How can you destroy creativity? Easy:

Chase success. Paul Graham says the way to destroy your creativity is to make money the top idea in your mind. It tends to consume all your mental energies. Blow also warns about thinking about success or how others will judge what you do. These things can easily lead to fear, and as Cleese says you need to feel confident to be able to generate ideas.

Obsess about disputes. Paul Graham talks about how Isaac Newton got involved in disputes and regretted the wasted energy. This is really just another form of worrying about what other people think.

Make a schedule. Blow warns about making a schedule, but also admits that we must all deal with schedules. Rich says his techniques don't work under pressure. While Cleese sets aside time to be creative, he recognizes that the process is unpredictable and needs time.

Pre-judge ideas. You must be open, Cleese doesn't call it "open mode" for nothing. Brainstorming forbids judging ideas, and as a technique it gets that much right. George Land found the more we use divergent and convergent thinking together—in other words the more we try to pre-judge ideas—the less creative we will be.

Get distracted. Cleese says you need to create a space free from distractions. For Blow, even the threat of a distraction can prevent him from relaxing, so he'll even spend time at a coffee shop for a few hours before heading into the office.

No humor. According to Cleese humor is about two frameworks coming together to make new meaning, and this is also the core of creativity. If you eliminate humor, then you eliminate creativity.

Live actively and urgently. If you want to ensure no relaxation happens, if you want to ensure that you are in closed mode, then live urgently and actively.

Conclusion


If you're here you are probably a computer programmer (most likely a Clojure programmer). That means you're probably a bit like me. You're good at thinking analytically and logically. You're good a judging solutions based on correctness, performance, etc. You're good at operating in "closed mode." These are great skills, and as Cleese says we need both open and closed mode to succeed, open to generate ideas, and closed to execute on them. We just may need to work on the open mode a bit.

You have the ability to be creative. You have a relaxed, curious, playful, imaginative self. There are some techniques that others have used that may help you access your creativity. They may help you, they may not. You may need to experiment a bit for yourself.

You cannot fully control this process. You can only indirectly stimulate creativity, and you cannot guarantee that your mind will solve the problem you want it to solve. One approach would be to work on several problems at once! You may also find some fruitful connections between the problems.

To be creative you must be persistent, and you must practice. I hope this helps you find those imaginative solutions.

Tuesday, August 30, 2016

"Clojure Polymorphism" Released!

From my new blog Real World Clojure. What am I doing with this new blog? I have no idea, but you can follow along.

~ ~ ~ ~

I have released a short e-book (30 pages) titled "Clojure Polymorphism." You can get 50% off by using this coupon link http://www.leanpub.com/clojurepolymorphism/c/ONeJZ629Isy7.
What is this book about?
When it comes to Clojure there are many tutorials, websites, and books about how to get started (language syntax, set up a project, configure your IDE, etc.). There are also many tutorials, websites, and books about how language features work (protocols, transducers, core.async). There are precious few tutorials, websites, and books about when and how to use Clojure's features.


This is a comparative architecture class. I assume you are familiar with Clojure and even a bit proficient at it.  I will pick a theme and talk about the tools Clojure provides in that theme.  I will use some example problems, solve them with different tools, and then pick them apart for what is good and what is bad.  There will not be one right answer.  There will be principles that apply in certain contexts.
I this installment, I will pick up the theme of "Polymorphism" looking at the tools of polymorphism that Clojure provides. Then I take a couple of problems and solve them several ways. At the end of it all, we look back at the implementations and extract principles. The end goal is for you to develop an understanding of tradeoffs and a taste for good Clojure design.


I have some ideas for other e-books. Perhaps a concurrency tour of Clojure taking a look at futures, STM, reducers, core.async, etc. Or maybe talk about identity by looking at atom, agent, ref, volatile!, etc. Or maybe look at code quality tools. Or how to organize namespaces. Or adding a new data structure with deftype?

What would you like to see? Contact me. :)

Friday, August 19, 2016

Reducible Streams

Laziness is a great tool, but there are some gotchas. The classic:

(with-open [f (io/reader (io/file some-file))]
  (line-seq f))

line-seq will return a lazy seq of lines read from some-file, but if the lazy seq escapes the dynamic extent of with-open, then you will get an exception:

IOException Stream closed  java.io.BufferedReader.ensureOpen (BufferedReader.java:115)

With laziness, the callee produces data, but the caller can control when data is produced. However, sometimes the data that is produced has associated resources that must be managed. Leaving the caller in control of when data is produced means the caller must know about and manage the related resources. Using a lazy sequence is like co-routines passing control back and forth between the caller and callee, but it only transfers control for each item, there is no way to run a cleanup routine after the caller has decided to stop consuming the sequence.

A Tempting Solution

One might immediately think about putting the resource control into the lazy seq:

(defn my-line-seq* [rdr [line & lines]]
  (if line
    (cons line (lazy-seq (my-line-seq* rdr lines)))
    (do (.close rdr)
        nil)))

(defn my-line-seq [some-file]
  (let [rdr (io/reader (io/file some-file))
        lines (line-seq rdr)]
    (my-line-seq* rdr lines)))

This way the caller can consume the sequence how it wants, but the callee remains in control of the resources. The problem with this approach is the caller is not guaranteed to fully consume the sequence, and unless the caller fully consumes the sequence the file reader will never get closed.

An Actual Solution

There is a way to fix this. You can require the caller to pass in a function to consume the generated data, then the callee can manage the resource and execute the function. It might look something like:

(defn process-the-file [some-file some-fn]
  (with-open [f (io/reader (io/file some-file))]
    (doall (some-fn (line-seq f)))))

(process-the-file my-file-name do-the-things)

Once upon a time clojure.java.jdbc used to have a with-query-results macro that would expose a lazy seq of query results, and you had these resource management issues. Then it was changed to use this second approach where you pass in functions.

There is a hitch to this approach. Now the callee has to know more about how the caller's logic works. For instance, in the above code you are assuming that some-fn returns a sequence that you can pass to doall, but what if some-fn reduces the sequence of lines down to a scalar value? Perhaps process-the-file could take two functions seq-fn and item-fn:

(defn process-the-file [some-file item-fn seq-fn]
  (with-open [f (io/reader (io/file some-file))]
    (seq-fn (map item-fn (line-seq f)))))

(process-the-file my-file-name do-a-thing identity)

That's better? I still see two problems:
  1. The caller is back to having to know/worry about resource management, because it could pass a seq-fn that does not fully realize the lazy seq before it escapes the with-open
  2. The logic hooks that process-the-file provides may never be quite right. What about a hook for when the file is open? How about when it is closed?
I could argue that this whole situation is worse, since the caller still has to worry about resource management, and now the callee has this additional burden of trying to predict all of the logic hooks the caller might want.

An additional design consequence is that you are inverting control from what it was in the lazy seq case. Whereas before the caller had control over when the data is consumed, now the callee does. You have to break your logic up into small chunks that can be passed into process-the-file, which can make the code a bit harder to follow, and you must put your sharded logic close to the callsite for process-the-file (i.e. you cannot take a lazy sequence from process-the-file and pass it to another part of your code for processing). There are advantages and disadvantages to this consequence, so it is not necessarily bad, it is just something you have to consider.

Another Solution

We can also solve this by using a different mechanism in Clojure: reduction. Normally you would think of the reduction process as taking a collection and producing a scalar value:

(defn process-the-file [some-file some-fn]
  (with-open [f (io/reader (io/file some-file))]
    (reduce (fn [a v] (conj a (somefn v)) [] (line-seq f))))

(process-the-file my-file-name do-a-thing)

While this may look very similar to our first attempt, we have some options for improving it. Ideally we'd like to push the resource management into the reduction process and pull the logic out. We can do this by reifying a couple of Clojure interfaces, and by taking advantage of transducers.

If we can wrap a stream in an object that is reducible, then it can manage its own resources. The reduction process puts the collection in control of how it is reduced, so it can clean up resources even in the case of early termination. When we also make use of transducers, we can keep our logic together as a single transformation pipeline, but pass the logic into the reduction process.

I have created a library called pjstadig/reducible-stream, which will create this wrapper object around a stream. There are several functions that will fuse an input stream, a decoding process, and resource management into an reducible object. Let's take a look at them:
  • decode-lines! will take an input stream and produce a reducible collection of the lines from that stream.
  • decode-edn! will take an input stream and produce a reducible collection of the objects read from that stream (using clojure.edn/read).
  • decode-clojure! will take an input stream and produce a reducible collection of the objects read from that stream (using clojure.core/read).
  • decode-transit! will take an input stream and produce a reducible collection of the objects read from that stream.
Finally, there is a decode! function that encapsulates the general abstraction, and can be used for some other kind of decoding process. Here is an example of the use of decode-lines!:

(into []
      (comp (filter (comp odd? count))
            (take-while (complement #(string/starts-with? % "1"))))
      (decode-lines! (io/input-stream (io/file "/etc/hosts"))))

This code will parse /etc/hosts into lines keeping only lines with an odd number of characters until it finds a line that starts with the number '1'. Whether the process consumes the entire file or not, the input stream will be closed.

Advantages:
  • This reducible object can be created and passed around to other bits of code until it is ready to be consumed.
  • When the object is consumed either partially or fully the related resources will be cleaned up.
  • Logic can be defined separately and in total (as a transducer), and can be applied to other sources like channels, collection, etc..
Disadvantages:
  • This object can only be consumed once. If you try to consume it again, you will get an exception because the stream is already closed.
  • If you treat this object like a sequence, it will fully consume the input stream and fully realize the decoded data in memory. In certain uses cases this may be an acceptable tradeoff for having the resources automatically managed for you.

Summary

Clojure affords you several different tools for deciding how to construct your logic and manage resources when you are processing collections. Laziness is one tool and it has advantages and disadvantages. It's main disadvantage is around managing resources.

By making use of transducers and the reduction process in a smart way, we can produce an object that can manage its own resources while also allowing collection processing logic to be defined externally. The library pjstadig/reducible-stream provides a way to construct these reducible wrappers with decoding and resource management fused to a stream.

Acknowledgments


Special hat tip to hiredman. His treatise on reducers is well worth the read. Many moons ago it got me started thinking about these things, and I think with transducers on the scene, the idea of a collection managing its own resources during reduction is even more interesting.

Tuesday, April 14, 2009

Terracotta Bug Reports

The integration work with Clojure and Terracotta has been progressing. After I brought up a couple of the issues I had encountered on the tc-dev mailing list, it turned out that I had discovered a couple of bugs in Terracotta.

These three issues were the source of some of the major changes that were required to integrate Clojure and Terracotta, and they should all be fixed in Terracotta 3.0.1:

In particular, CDV-1233 required some ugly changes in the compiler, which should now be (thankfully!) unnecessary.

I believe I also have a solution to the problem that comes about when a root Var binding is a non-portable object. Stay tuned for that!

Monday, March 30, 2009

Clojure + Terracotta Update

I have gotten to the point in my Clojure + Terracotta experiment, where I believe all of the features of Clojure are functional (Refs, Atoms, transactions, etc.). I do not have a way to extensively test the Clojure functionality, but I have run the clojure.contrib.test-clojure test suites successfully, as well as some simple tests on my machine.

There are still some open issues, and given the limited extent to which I have tested this, I would not consider this production quality in the least. I would welcome help from the Clojure community in testing this integration module. I'm sure there are unexplored corners.

Being that several of the changes are relatively trivial, they could be easily integrated back into the Clojure core. I have detailed as best as possible the changes I had to make to Clojure in this report: Clojure + Terracotta Integration Report

The code is available at GitHub (http://github.com/pjstadig/tim-clojure-1.0-snapshot/tree/master), and there are instructions on setting it up and running the code. If you have any difficulties or questions, please feel free to e-mail me paul@stadig.name

Thursday, March 5, 2009

Clojure + Terracotta: We Have REPLs!

Update: The Clojure + Terracotta integration is (I believe) feature complete. Details at http://paul.stadig.name/2009/03/clojure-terracotta-update.html.

JVM #1

paul@pstadig-laptop:~/tim-clojure/tim-clojure-1.0-SNAPSHOT/examples/shared-everything$ ./bin/dso-clojure repl.clj 
Starting BootJarTool...
2009-03-05 15:00:10,868 INFO - Terracotta 2.7.3, as of 20090129-100125 (Revision 11424 by cruise@su10mo5 from 2.7)
2009-03-05 15:00:11,428 INFO - Configuration loaded from the file at '/home/paul/tim-clojure/tim-clojure-1.0-SNAPSHOT/examples/shared-everything/tc-config.xml'.

Starting Terracotta client...
2009-03-05 15:00:14,904 INFO - Terracotta 2.7.3, as of 20090129-100125 (Revision 11424 by cruise@su10mo5 from 2.7)
2009-03-05 15:00:15,436 INFO - Configuration loaded from the file at '/home/paul/tim-clojure/tim-clojure-1.0-SNAPSHOT/examples/shared-everything/tc-config.xml'.
2009-03-05 15:00:15,656 INFO - Log file: '/home/paul/terracotta/client-logs/org.terracotta.modules.sample/20090305150015636/terracotta-client.log'.
2009-03-05 15:00:17,870 INFO - Statistics buffer: '/home/paul/tim-clojure/tim-clojure-1.0-SNAPSHOT/examples/shared-everything/statistics-127.0.1.1'.
2009-03-05 15:00:18,421 INFO - Connection successfully established to server at 127.0.1.1:9510
user=> (defn foo [] 42)
#'user/foo
user=>

JVM #2

paul@pstadig-laptop:~/tim-clojure/tim-clojure-1.0-SNAPSHOT/examples/shared-everything$ ./bin/dso-clojure repl.clj
Starting BootJarTool...
2009-03-05 15:01:39,663 INFO - Terracotta 2.7.3, as of 20090129-100125 (Revision 11424 by cruise@su10mo5 from 2.7)
2009-03-05 15:01:40,225 INFO - Configuration loaded from the file at '/home/paul/tim-clojure/tim-clojure-1.0-SNAPSHOT/examples/shared-everything/tc-config.xml'.

Starting Terracotta client...
2009-03-05 15:01:45,507 INFO - Terracotta 2.7.3, as of 20090129-100125 (Revision 11424 by cruise@su10mo5 from 2.7)
2009-03-05 15:01:46,091 INFO - Configuration loaded from the file at '/home/paul/tim-clojure/tim-clojure-1.0-SNAPSHOT/examples/shared-everything/tc-config.xml'.
2009-03-05 15:01:46,275 INFO - Log file: '/home/paul/terracotta/client-logs/org.terracotta.modules.sample/20090305150146254/terracotta-client.log'.
2009-03-05 15:01:50,868 INFO - Connection successfully established to server at 127.0.1.1:9510
"user=> "(foo)
42
"user=> "*print-dup*
false
"user=> "

Commentary

This is obviously an example of the "shared everything" approach. It's neither perfect nor complete, but it's a start. There are still some non-portable classes that need to be reworked, for some reason the second VM is printing the REPL prompt readably even though *print-dup* (as you can see) is false, I still haven't worked out the problem with *in*, *out*, and *err*, etc. etc.

It's still very raw, but I'll see if I can't push to github in the next day or two. This is an exciting first step!

Tuesday, March 3, 2009

Clojure + Terracotta: The Next Steps

Update: I've gotten a *multiple* REPLs running with Terracotta. http://paul.stadig.name/2009/03/clojure-terracotta-we-have-repl.html.

In my last post about Clojure + Terracotta I gave an example of sharing specific references between JVMs through Terracotta. This is what I call the "shared somethings" approach. You specify exactly what you like to share. Another approach is what I call the "shared everything" approach.

Shared Everything

The goal of shared everything is to have multiple VMs working within a single global context through Terracotta, all of your vars and refs would be shared by default, and the canonical test case for this would be to define a function in one VM and have it show up automatically in another VM.

The first task was to move my work into a Terracotta Integration Module (TIM). When using a TIM, in addition to packaging the configuration for reuse, classes can be replaced with clustered versions that will work with Terracotta, without having to fork the original code base.

The second was to replace a couple of classes in the Clojure codebase. The Namespace class uses AtomicReference, which is not supported by Terracotta. There was a minor change necessary in Var, too, because it was using its dvals field as a sentinel value to indicate that the var is not bound. This does not do for Terracotta, because dvals is a ThreadLocal, so I created a NOT_BOUND sentinel field. There were some other changes as well, I'm not going to detail all of the changes, but you get the idea.

At this point I would have hoped that I could run a REPL and possibly even try my canonical test case, but it should be so easy. I have run into two major roadblocks:

  1. *in*, *out*, and *err*. Being I/O streams, *in*, *out*, and *err* obviously cannot be shared through Terracotta. The problem is that they are stored as Vars and interned into the clojure.core namespace. This means that Terracotta will try to share them, because they are part of the shared object graph. I could make clojure.lang.Var.root a transient field (through Terracotta's configuration file), but that would make the roots of all Vars transient, which is not what we want. Instead, what I thought I needed is some kind of TransientVar that could have a different root value (not just bindings) for each JVM. I pursued this a bit using the class replacement of the TIM, and concluded that if that is the route to go, then it should probably be made in the Clojure code (it got messy), or that at least there are some changes to the Clojure code that would ease this. What I settled on (after a suggestion from Rich) is to leave the root bindings for *in*, *out*, and *err* as nil and allow the REPL to bind them. However, the REPL did not bind them for me, so I created my own repl.clj file that binds them and calls clojure.main/repl, and it works! However, this is only a temporary solution. Whether it is creating a TransientVar class, or something else, we need a more permanent solution.
  2. Classes. I am able to connect a single JVM to Terracotta, and run the REPL, but I cannot connect multiple JVMs, nor can I disconnect and reconnect a single JVM. When an instance of Clojure connects to Terracotta, it pulls a compiled function out of the object cache, and then throws a ClassNotFoundException because it cannot find the associated class. I started to pursue modifying the DynamicClassLoader and Compiler to store the compiled classes in the Terracotta object graph, and I still think that this might be the direction to go in. However, I wanted to go ahead and share what I have and get some feedback to see if there are any other solutions.

The code is available at http://github.com/pjstadig/tim-clojure-1.0-snapshot/tree/master. In the "examples" directory I have a "shared-everything" example and a "shared-somethings" example. If you have any trouble running the examples, then let me know. There are some dead ends and some commented code that may not make sense, but my goal was to do a proof-of-concept first and to clean it up once I understand what needs to be done.

Conclusion

We are getting close to a "shared-everything" approach to integrating Clojure and Terracotta. We have some issues to deal with, but we are on our way to making this dream a reality.

Friday, February 27, 2009

Clojure + Terracotta = Yeah, Baby!

Update: I've gotten a REPL running with Terracotta. http://paul.stadig.name/2009/03/clojure-terracotta-next-steps.html.

What is Terracotta?

Terracotta provides a network-attached, virtual, persistent heap and transparent inter-JVM thread coordination. With Terracotta, you no longer need to map your objects to database tables and back. You simply hand your object to Terracotta and it will cache your data. Not only does it cache your data, but it will make your object available to a cluster of networked JVMs. Not only that, but it will also spill your objects to disk if necessary (just like Virtual Memory), so you need not worry about having gobs of memory to hold all of your objects.

What is Clojure?

Clojure is a Lisp for the JVM with a software transactional memory, and agents (asynchronous, message based concurrency). It is a functional language with immutable datatypes. It can also inter-operate with any existing Java code.

NOTE: you need to use Clojure r1310 or later, because the Keyword class needs to have hashCode implemented to play nicely with Terracotta.

Clojure + Terracotta = ?

These two seem like an interesting combination. Imagine the possibilities...kill your database, simple POJO applications, free distributed transactions, clustered JVMs with limitless memory...it would make your hair would grow back, you'd get women, and become filthy rich...well...maybe not, but at least you'd have more fun writing software.

After some initial setup (the code and instructions are at http://github.com/pjstadig/terraclojure/tree/master/), there are two things that need to be done to integrate Clojure and Terracotta: 1) instead of running Clojure with the 'java' command, you run it with the 'dso-java.{sh,bat}' script provided with Terracotta, and 2) you need to create a configuration file that defines how your objects will be shared between JVMs.

Configuration

The configuration for Terracotta (at least in our case) consists of defining: roots, instrumented classes, auto-locks, additional boot jar classes, and servers. At this point it's probably helpful to take a peek at the config.xml file that comes with the code and follow along.

  • Roots. A root is a object that is shared between JVMs. Any objects that are part of the object graph that can be reached from the root are also shared, so any objects that are assigned to data members, etc. A common use case is to have a ConcurrentHashMap (or in our case a PersistentHashMap from Clojure) that is shared as a root. This creates a flexible hierarchy of shared objects. In Clojure's case, we also share clojure.lang.Keyword.table, so that our keywords are unique across all of the JVMs, otherwise inserting into a hash map would create multiple entries for the same keyword.
  • Instrumented classes. Any class that is shared (either directly as a root, or indirectly as a part of a root's object graph), must be instrumented. I made all of the clojure.lang.* classes instrumented. It's a bit of a broad stroke, but there aren't any performance problems that result from instrumenting too many classes. Terracotta is helpful in this case, if you end up inserting an uninstrumented class into the object graph, it'll throw a RuntimeException that explains exactly how to modify your config file to instrument that class.
  • Auto-locks. Terracotta will transparently convert your synchronized blocks into distributed transactions across all the JVMs in the cluster. Again, I made broad strokes here and just defined auto-locks for all of the methods on any clojure.lang.* class, and again, there aren't any performance penalties for auto-locking methods that don't have any synchronized blocks. I used write locks, and Terracotta has a few different types of locks that are worth looking into if you need to do something more serious. In the case of auto-locks, Terracotta will also help you out by throwing a RuntimeException if you leave out anything.
  • Additional boot jar classes. Frankly, this was something Terracotta told me to do, and I don't know exactly what is going on here. (Perhaps someone else can explain?) I think what happens is that by default Terracotta instruments the java.lang.* and java.util.concurrent.* classes, but to instrument other Java core classes you have to add them in this configuration element.
  • Servers. Terracotta is very easy to work with, and by default will just run a single server on localhost. You can define more than one server in a cluster. In my case, I only wanted one server, but I wanted to change the persistence mode. By default the persistence mode is a temporary-swap-only mode. The objects will be preserved across stopping and starting clients, but once the server is stopped, the data disappears. To have the objects persisted across restarting the server, you have to set the persistence mode to permanent-store. The temporary swap mode will be faster for data like the intermediate results of calculations, caching, etc., but if you need to permanently persist the data, then you need to use permament-store.

There are instructions about how to run this example in the README with the code, so I won't bother to duplicate that here. I'd just like to share some of the issues I encountered, the results, and any future direction that could be taken.

Issues

The first major issue that I encountered was that Keywords weren't unique across JVMs, so I had to make clojure.lang.Keyword.table a root. This ensured that keywords are unique across JVMs, but I still ran into an issue when using keywords as keys for a PersistentHashMap. The result of identical? was true for keywords from two JVMs, but I was still getting duplicate entries in my hash map. After some debugging, I was able to determine that the issue is that the keyword class did not override the default implementation of hashCode. After mentioning this to Rich, and a quick fix in r1310, it worked nicely.

The only other major issue was how to reference Clojure vars and refs from the Java side. The main reason for this is to define a root that will be shared by Terracotta. When Clojure code gets compiled some Java classes get generated with mangled names. As far as I can tell, there isn't a good predictable way to get at a Clojure var, because Clojure will generate a class for each namespace called my/namespace/namespace__init.class and it creates static fields on that class for various definitions (functions, vars, etc.). Those fields are named const_1, const_2, const_3, etc. There is no reliable, flexible way to predict the name of a particular Var.

My solution was to create a simple Java class called terraclojure.Root with a couple of static fields containing refs. At first I just used that class directly to access the refs, but then I decided to actually assign the static fields to some vars in my namespace, i.e. (def *hash* terraclojure.Root/hash). This works and it makes it a little more transparent on the Clojure side. I would be happy to hear if there is another way to do this.

Result

The result of this whole experiment was that I am able to use the Software Transactional Memory with a couple of refs, and to have my changes shared across multiple JVMs. I didn't do any extensive testing to verify that transaction retries work as expected, but since Clojure uses the java.util.concurrent.* classes and standard synchronization, I don't expect there would be an issues.

Where do we go from here?

I only experimented with the STM. I didn't experiment with Agents, so that is certainly an area for future work. On the Terracotta side, I only used one server, I didn't setup a whole array of servers, nor did I try using one or more servers on different machines. All my testing was local, so the performance reflected that (it was pretty good! :)). If you do any further experimentation, then please share it on a blog or to the Google group.

Conclusion

I don't have a lot of experience with Terracotta, but it seems to be quite mature and easy-to-use. I also think that Clojure is a very exciting language, and the combination of the two opens up some interesting possibilities for how to architect highly available, scalable, database-less applications.

P.S. I have a B.S. in Computer Science and will have an M.S. in Computer Science in May. I don't do anything near this interesting at my job. If you have any need for consulting, or if you'd like to offer me a job ;), then feel free to contact me at paul@stadig.name.

Tuesday, November 4, 2008

Clojure: a LISP that has a chance

Clojure is an interesting new language. Here's the executive summary:

Clojure is a dynamic programming language that targets the Java Virtual Machine. It is designed to be a general-purpose language, combining the approachability and interactive development of a scripting language with an efficient and robust infrastructure for multithreaded programming. Clojure is a compiled language - it compiles directly to JVM bytecode, yet remains completely dynamic. Every feature supported by Clojure is supported at runtime. Clojure provides easy access to the Java frameworks, with optional type hints and type inference, to ensure that calls to Java can avoid reflection.

Clojure is a dialect of Lisp, and shares with Lisp the code-as-data philosophy and a powerful macro system. Clojure is predominantly a functional programming language, and features a rich set of immutable, persistent data structures. When mutable state is needed, Clojure offers a software transactional memory system and reactive Agent system that ensure clean, correct, multithreaded designs. — www.clojure.org

Each of the key features is exciting to me, a functional LISP that integrates closely with the JVM and has baked-in concurrency. What's more, I think Clojure has a chance at making it big. Here's why:

  • Unique Vision. I don't think any new language can survive for long, unless it has a unique vision. Clojure's unique vision is to bring together a mix of performant, immutable data stuctures, baked-in concurrency, functional style, and close integration with the JVM.
  • JVM Integration. Rich Hickey had the incredible foresight to see the JVM as a platform to be embraced closely. This means that not only can you leverage 100% of existing Java code, but your Clojure code compiles to Java bytecode and benefits from the HotSpot JVM's dynamic optimizations. Compare that to a "from scratch" language that takes years to get a diverse set of libraries and an optimized implementation.
  • Benevolent Dictator. I have always thought that a new LISP (or any new language for that matter) needs a Benevolent Dictator. The BD is the friendly face of the community and sets the tone for how people treat each other. But more importantly the BD is a dictator who has a strong vision for the language, and will say "no" to feature requests that don't line up with his vision. This is Rich Hickey, friendly and open to suggestion, but not afraid to say "no."

If any of this sounds interesting to you, then check out the homepage. The quickest and easiest way to get involved in the community is to join the Google Group. Also, if you're into IRC, then check out #clojure on irc.freenode.net.

I am excited about the future of Clojure, and have really enjoyed working with it so far.