{"id":35,"date":"2010-09-26T01:58:08","date_gmt":"2010-09-26T01:58:08","guid":{"rendered":"http:\/\/blogs.sun.com\/clayb\/entry\/when_do_people_work"},"modified":"2021-02-16T18:06:46","modified_gmt":"2021-02-16T18:06:46","slug":"when-do-people-work","status":"publish","type":"post","link":"https:\/\/clayb.net\/blog\/when-do-people-work\/","title":{"rendered":"When do people work?"},"content":{"rendered":"<h2>Ever wonder when people are actually working?<\/h2>\n<p>It can be hard answering, &#8220;when are people at work?&#8221; On a distributed team, with many co-workers and the typical corporate dotted-line type relationships, it is even harder! Inevitably communications on schedule shifts and desired schedules go un-communicated. A few years ago, this occurred for folks I worked with.<br \/>\n<!--more--><\/p>\n<h2>So what does this mean?<\/h2>\n<p>I decided to graph people&#8217;s e-mail times, as below:<\/p>\n<p><center><a href=\"https:\/\/clayb.net\/blog\/wp-content\/uploads\/2012\/01\/publishable_hists.png\"><img decoding=\"async\" src=\"https:\/\/clayb.net\/blog\/wp-content\/uploads\/2012\/01\/publishable_hists.png\" alt=\"\" width=\"75%\" height=\"75%\"><\/a><\/center>This gives a good perspective of when folks are in the office (or at least productive and sending e-mails). For example, here we have that one coworker in Colorado is actually closer to the times worked by folks in Europe and a large portion of the team &#8211; independent of location &#8211; work centered around noon mountain time; however, there are those of us (myself included) who trend later. Also visible is an unfortunate effect of the chosen statistical method; times across midnight are not properly accounted for resulting in outliers for folks who do trend later and send off e-mails around 1 or 2am while a more robust technique would likely take these into better account and show a later median for these coworkers.<\/p>\n<h2>What catalogs when we actually work?<\/h2>\n<p>To try and passively monitor working hours was my desire. I had assumed actively asking would be biased or again go forgotten when schedules shift (e.g. school year starts for one&#8217;s kids or summer leads to desired morning bicycle commutes). Further, actively asking people to &#8220;punch-in&#8221; and &#8220;punch-out&#8221; would be a pain and pretty foreign to engineers. So, that begs the question, what is a good proxy of working hours?<\/p>\n<p>For many teams, <a href=\"http:\/\/en.wikipedia.org\/wiki\/Internet_Relay_Chat\">IRC<\/a> or instant messenger log-in and log-off times can work. But of course, some folks stay logged-in 24 hours by 7 days a week; or drop off often due to flaky network connections. Code commit and bug filing times can work, if everyone on the team is doing a number of code commits and bug files. However, if these events are relatively rare it makes the proxy less valuable. For my research, I settled on e-mail times as e-mail is very popular in communities I work in.<\/p>\n<p>Now, unless you happen to be a pack rat, it can be difficult to muster a large corpus of e-mail data but luckily, if you work on an open source project &#8212; with an external mailing list &#8212; then all data is retained in the vast archives of usually a <a href=\"http:\/\/www.gnu.org\/software\/mailman\/index.html\">GNU Mailman<\/a> list. There are various niceties about Mailman, one can use the files retained on the server or simply trawl the <a href=\"http:\/\/mail.opensolaris.org\/pipermail\/\">Pipermail<\/a>web interface for data on who posted to the list when. Then it&#8217;s a simple matter to develop an XML or CSV (comma separated value) file of who posted what when and use your favorite graphing package to view the data.<\/p>\n<p>I choose to use Python with <a href=\"http:\/\/codespeak.net\/lxml\/\">LXML<\/a> to parse the OpenSolaris Mailman archives for the list I was interested in and produce an XML representation of the data (see the script <a href=\"https:\/\/clayb.net\/blog\/wp-content\/uploads\/2012\/02\/downloadPipermail.py\">here<\/a>). From this, I was able to easily construct a CSV file which I could load into <a href=\"http:\/\/www.r-project.org\/\">GNU R<\/a> for slicing in interesting ways.<\/p>\n<h2>Great, data&#8217;s fun and all &#8212; but now what?<\/h2>\n<p>In my case, I was mostly interested in co-workers around North America to see my relative working hours and averages. (I was looking at a pretty short period of time so unconcerned with seasonality and other variance.)<\/p>\n<p>Especially in engineering, working hours can range from 40-50 hours a week to &#8220;crunch time&#8221; sprints of 60-70+. As such, how can one try to normalize out and see trends in data so that a 4am &#8220;crunch time&#8221; e-mail does not throw off an average otherwise around 9am-5pm? Though, not quite the most robust technique I settled on a series of box plots for my coworkers. However, this came with pitfalls for simply comparing those who would be working in very distant timezones.<\/p>\n<p>Using a simple time format of 0-2359 for 12:00am to 11:59pm, I viewed everyone&#8217;s e-mail posts in my local timezone. Unfortunately, this is not so robust as someone in Europe will span from around 8pm to 8am which in my representation will be a boxplot with a median of noon my time but in reality is the inverse of what is shown. (Notice that the folks in California in my example image have a lot of &#8220;outliers&#8221; in the morning.) However, being roughly in the temporal middle of the Americas living in Colorado, this approach worked reasonably well for comparing with other Americans. To finally view the data I simply used a GNU R <a href=\"https:\/\/clayb.net\/blog\/wp-content\/uploads\/2012\/02\/publishable_times.r\">script<\/a> and a hastily written script to chomp down some CSV <a href=\"https:\/\/clayb.net\/blog\/wp-content\/uploads\/2012\/02\/publishable.csv\">data<\/a> I produced from the <a href=\"https:\/\/clayb.net\/blog-aux\/publishable.xml\">XML<\/a> I had scraped.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Ever wonder when people are actually working? It can be hard answering, &#8220;when are people at work?&#8221; On a distributed team, with many co-workers and the typical corporate dotted-line type relationships, it is even harder! Inevitably communications on schedule shifts and desired schedules go un-communicated. A few years ago, this occurred for folks I worked [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6,7,8,26,10],"tags":[],"_links":{"self":[{"href":"https:\/\/clayb.net\/blog\/wp-json\/wp\/v2\/posts\/35"}],"collection":[{"href":"https:\/\/clayb.net\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/clayb.net\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/clayb.net\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/clayb.net\/blog\/wp-json\/wp\/v2\/comments?post=35"}],"version-history":[{"count":0,"href":"https:\/\/clayb.net\/blog\/wp-json\/wp\/v2\/posts\/35\/revisions"}],"wp:attachment":[{"href":"https:\/\/clayb.net\/blog\/wp-json\/wp\/v2\/media?parent=35"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/clayb.net\/blog\/wp-json\/wp\/v2\/categories?post=35"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/clayb.net\/blog\/wp-json\/wp\/v2\/tags?post=35"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}