Monday, March 24, 2014

The Three Co-ordinators

It is has been a while since we posted on the blog. Generally, this means that things have been busy and interesting. Things have been busy and interesting.

We are presently, going through redevelopment of the site, the evaluation of new techniques for service delivery such as using Docker for containers and updating multiple services throughout the sites.

The development of the programme presented at CHEP on automation and different approaches to delivering HEP related Grid services is underway. An evaluation of container based solutions for service deployment will be presented at the next GridPP collaboration meeting later this month. Other evaluation work on using Software Defined Networking hasn't progressed as quickly as we would have like but is still underway.

Graeme (left), Mark (center) and Gareth.

On other news, Gareth Roy is taking over as the Scotgrid Technical Co-ordinator this month. Mark is off for adventures with the Urban Studies Big Data Group within Glasgow University.And as Dr Who can do it, we can do. Co-ordinator Past, Present and Future all appear in the same place at the same time.

Will the fabric of Scotgrid be the same again?

Very much so.

Monday, October 14, 2013

Welcome to CHEP 2013

Greetings from CHEP 2013 in a rather wet Amsterdam.

The conference season is upon us and Sam, Andy, Wahid and myself find ourselves in Amsterdam for CHEP 2013. CHEP started here in 1983 and it is hard to believe that it has been 18 months since New York.

As usual the agenda for the next 5 days is packed. Some of the highlights so far have included advanced facility monitoring, the future of C++ and Robert Lupton's excellent talk on software engineering for Science.

As with all of my visits to Amsterdam, the rain is worth mentioning. So much so that it made local news this morning. However, the venue is the rather splendid Beurs van Berlage in central Amsterdam.

CHEP 2013


There will be further updates during the week as the conference progresses.




Busy Year

We haven't posted a great deal this year as there has been a huge amount going on within Scotgrid since January.

The main news this year has been Stuart's departure to Saint Andrews University from the Glasgow Scotgrid Team. Stuart's role within EGI, ROD and Grid Ops as well as the Glasgow site and his development work on MPI at Glasgow says a lot about his rather busy part-time role within Scotgrid. The word Factotum or "make everything" springs to mind when describing his input.
We all wish Stuart the very best at Saint Andrews.

The Glasgow site has suffered from issues with the coolant infrastructure since January. To mitigate this the University is upgrading both the power and air conditioning within the Kelvin Building. This work will include the installation of a Generator and UPS system as well as new air conditioning units. This is a long term project and will be completed by the summer of 2014.

ECDF has performed incredibly well since January and Durham, while suffering from air con and power issues earlier in the year is now relatively stable.

We have brought in additional VOs with the MVLS group at Glasgow University and are presently in discussions with other non-HEP groups such as bio-chemistry. The most technically challenging project is the proposed investigation into the Lairg Magnetic anomaly by the EarthSci group at Glasgow. This project is difficult due to the lack of network connectivity in the area where the data is being generated, we will report on this soon.

Our primary focus of research, outside of running the sites, have covered GPU work at ECDF, more efficient data management and deployment strategies at Glasgow and ECDF. Additionally, how we utilise containerisation and build smarter cluster restart environments has been investigated by Gareth at Glasgow. David Crook's has done excellent work around aggregating the multiple monitoring platforms that have sprung up within the Grid by utilising the Graphite package.

We have attended and presented in multiple conferences and public outreach events including one during the Edinburgh Festival.

So that is up to date in time for CHEP. Which is on this week. Still trying to work out how quickly the last 18 months went.


Tuesday, December 25, 2012

A Merry Christmas and a Happy New Year to all our followers, users and co-workers from Scotgrid Glasgow.
Keeping to a Physics theme, as always.

See you all in 2013.

Thursday, December 06, 2012

2012; A Grid Odyessy

We haven't published on the blog since September this year which is a bit remiss of us.
There are many reasons for this. Primarily we have been working through the final back log of the DRI grant until October. The expansion of the Glasgow site to 4000 cores, terra scale networking and changes to the disk farm have not been simple. Once the long standing issues with the internal data network were resolved with the upgrade to the Extreme Networks equipment, additional issues around the placement of data by DPM became evident. This was not a non trivial task to investigate. Stuart and Sam are in the process of developing a software patch for allowing a more sensible placement of data files within the cluster.

In addition to this work, we are currently considering software and hardware changes to our data storage architecture in the new year. More of this in January.

Again this year Glasgow has been plagued with infrastructure issues which have caused several major issues to the site's operation. We are now in a position where there is a major upgrade programme underway to deliver a more robust power, fire suppression and air conditioning system throughout the computer rooms.

While these combined issues have caused a large number of issues the Glasgow site saw a return to 100 % availability and reliability metrics for November on the WLCG accounting earlier this week.
Hopefully, this is how we will continue through the Christmas period and into 2013.

As the end of winter is upon us with the winter solstice being just over 15 days away and Christmas following shortly behind it we would like to wish everyone a Merry Christmas and a Happy New Year for all of those at Scotgrid.

Monday, September 03, 2012

All hands to the pumps, oh wait, that is Glycol

Unfortunately on Friday the Air Con fairy visited Glasgow and due to a faulty pressure valve decided to sprinkle some magic in one of our plant rooms by dumping liquid coolant onto the floor. We took emergency action and closed down the cluster as the heat being generated in 141 was going above 30 degrees centigrade. The faulty equipment and associated devices were replaced. Thankfully, there hasn't been any damage to the equipment and we will be coming out of downtime and going back into production shortly.

The Higgs appears

This year we haven't been great at keeping up with blog posts but there are many reasons behind this. We have installed a new network, additional cores taking us up to 4000 available job slots and have upgraded the server infrastructure throughout  GU Scotgrid's cluster. Also, we have fallen foul of infrastructure issues and have had problems with the old and replacement Air Con systems. Slowly, we are extracting ourselves from these issues and  recently Professor David Britton gave a lecture on the Grids role in the announcement made in July of this year at CERN during this years Turing Festival. A surprise appearance at the event was Professor Higgs himself.

Professors David Britton (left) and Professor Peter Higgs

Also presenting at the event were Professors Tejinder Singh Virdee and John Ellis of Imperial College and Dr Ben Segal from CERN. The event was one of the kickstart activites for the Turing Festival and enabled the public and academics to get a better over view of what has been involved in getting the experiments this far.


Friday, August 03, 2012

Scotgrid Calling

It has been a while since we last updated the blog which is generally a sign of being busy. Unfortunately we have encountered several infrastructure issues recently which needed to be repaired. Predominantly these revolved around the air conditioning units on the roof of the Kelvin Building. This work was completed a few weeks ago but as one thing has been fixed another issue identified itself in the form of a failing Air Handling Unit in 141. The knock on effect of this is that we can't take full advantage of the cluster servers located in the room and the overall cluster is running at two thirds capacity presently.

While, these events are less than optimal it has allowed us to plan the next set of cluster upgrades which will introduced another 256 job slots into the cluster and due to the new resilient network fabric we have developed the deployment of these services is no longer limited to one room supporting 10 gig interfaces.

Other developments also include the re-introduction of an independent control network and a new WAN testing platform Perfsonar. We will blog about this seperately shortly.


Thursday, May 24, 2012

CHEP Update

As we are into the 4th Day of CHEP a quick overview of the activities of Scotgrid, GridPP and the conference as a whole is now in order.

We have presented our posters and have generated interest in Storage, Job failures, Network Security and some of the work we have conducted with IPv6. Several potential collaborations with other sites and developers have resulted from these presentations. Andy and Wahid had several successful talks and there was a high volume of interest on the work being discussed.

From a GridPP perspective Chris Walker's poster on using Lustre for low cost petascale storage also generated a large volume of interest. Talks given by other members of the collaboration were equally well received.

The conference itself has covered the multiple developments within the field over the last 12 - 18 months with presentations investigating a variety of topics including federations for data, the future of CPUs/GPU, ultra high speed networking and common software architectures for the experiments.

The variety of techniques being deployed and approaches taken to Grid centric problems are always of interest.


Saturday, May 19, 2012

Scotgrid in the Big Apple for CHEP


We are attending the WLCG and CHEP in New York this week. There will be regular updating of the blog with details of the talks and papers we are attending.

Monday, May 07, 2012

Stockholm LHCONE Meeting

Kristall in the Sergels Torg Stockholm


We were in attendance at the LHCONE  meeting at KTH in Stockholm last week. The purpose of this collaboration is to investigate the efficient use of networks globally for LHC research. As usual it was an excellent meeting where the technical mechanisms for current and future network deployments were discussed and considered.

The agenda can be found here. Some of the highlights of the meeting included an excellent presentation by Erwin Laure on the Swedish and Scandinavian Super Computing and Grid computing infrastructure, Joe Mambretti's presentation on the GLORIAD global research network, Mike O'Connor's discussion on the technical configurations required to avoid asymmetric routing issues between the LHCONE and the current production networks and  Domenico Vicinanza's presentation on Perfsonar MDM.

In addition to these presentations technical discussions surrounding various technologies surrounding bandwidth reservation, ultra high speed networking and  Open Flow technologies were held. As these discussions develop through the network architecture groups we will keep you up to date. 


Also, the weather in Stockholm was exceptional and the KTH Campus is worth a visit for its architecture alone. I would like to thank our hosts and all the other attendees for making this such an enjoyable and informative couple of days.




 KTH Campus Stockholm
 

Sunday, May 06, 2012

Preparing for IPv6

Generally, we don't repost news items on the blog but this BBC article gives a good indication of the changes underway globally for implementing IPv6. currently the Glasgow Scotgrid test cluster is being revamped post our last spending cycle and we are embarking on a full test programme of IPv6 specifically around running Grid services. As this work progresses we will regularly update the blog.

Thursday, May 03, 2012

GridPP At The Top Of Europe

This news article appeared on the GridPP website and is worth reposting to our blog as it gives an overview of the collaborations efforts to date within the WLCG and with the Non High Energy Physics (HEP) communities.
GridPP At The Top Of Europe

Tuesday, April 17, 2012

One XOS, Great Big Purple Packet Eater. Sure looks good to me.

So we haven't been blogging a great deal since December and for good reason. We found ourselves in the exciting position of being given additional funding to enhance our network capability and also we had additional equipment to install into the cluster.

First things first however, as you may have read we have had no end of issues with the older network equipment. We had a multi-vendor environment which, while adequate for 800 analysis jobs and 1200 production jobs, wasn't quite up to cutting the mustard as we couldn't expand from there.

The main reason was the 20 Gig link between the two computing rooms which was having real capacity issues. Also, add in issues between the Dell and Nortel LAG and associated back flow problems, sprinkled with a buffer memory issue on the 5510s and you get the picture. In addition to this we were running out of 10 Gig ports and therefore couldn't get much bigger without some investment.

Therefore, the grant award was a welcome attempt to fix this issue. After going to tender we decided upon equipment from Extreme Networks. The proposed solution allowed for a vast 160 Gigabit interconnect between the rooms broken into two resilient link bundles in the Core and an 80 Gigabit Edge layer. In addition to this connection we also installed a 32 core OM4 grade fiber optic network for the cluster which will carry us into the realms of 100 Gigabit connections, when it becomes available and cheap enough to deploy sensibly.

We now have 40 x 40 gigabit port, 208 x 10 gigabit ports and 576 1 x Gigabit ports available for the Cluster.

 There is quick and clever and here it is

The new deployment utilises X670s  in the Core and X460s at the Edge.

The magic of the new Extreme Network is that it uses EAPS, so bye bye Spanning Tree and good riddance as well as MLAG which allows us to load share traffic across the two rooms so having 10 Gigabit connections for disk servers in one room is no longer an issue.

Then it got a bit better. Due to the Extreme OS we can now write scripts to handle events within the network which ties in with the longer term plan for a Cluster Expert System (ARCTURUS) which we are currently designing for test deployment. More on this after August.

Finally, it even comes with its own event monitoring software, Ridgeline which gives a GUI interface to the whole deployment.

We stripped out the old network installed the new one and after some initial problems with the configuration, which were fixed in a most awesome fashion by Extreme got the new one up and running. What we can say is that the network isn't a problem anymore, at all.

This has allowed us to start to concentrate upon other issues within the Cluster and look at the finalised deployment of the IPV6 test cluster which has benefited in terms of hardware from the new network install. Again, more on this soon.

Right, so now to the rest of the upgrade we have also extended our cold isle enclosure to 12 racks, have a secondary 10 Gig link onto the Campus being installed and have a UPS. In Addition to this we refreshed our storage using Dell R510s and M1200s as well as buying 5 Interlagos boxes to augment the worker node deployment.

 The TARDIS just keeps growing

We also invested in an experimental user access system with wi-fi and will be trying this out in the test cluster to see if a wi-fi mesh environment can support a limited number of grid jobs.  As you do.

In addition to this we improved connectivity for the research community in PPE at Glasgow and across the Campus as a whole, with part of the award being used to deliver the resilient second link and associated switching fabrics.

It hasn't been the most straight forward process as the decommissioning and deployment work was complex and very time consuming in an attempt to keep the cluster up and running as long as possible and to minimise down times.

We didn't quite manage this as well as expected due to the configuration issues on the new network but we have now upgraded the entire network and have removed multiple older servers from the cluster to allow us to enhance the entire batch system for the next 24 - 48 months.

As we continue to implement additional upgrades to the cluster we will keep you informed.
For now it is back to the computer rooms.

Monday, February 27, 2012

LSC files and emailAddress redux

This post involves a very complicated journey to get to a simple place.

The fundamental problem is around the catchy titled OID 1.2.840.113549.1.9.1

No, wait, let me take a step back. On the Grid, we use certificates for authentication. An X509 certificate is, as with most certificates, a signed set of assertions, and a public key. As with the rest of the X500 standards, it's native language is something called ASN.1 (Abstract Syntax Notation 1) (aka X208, and the later revision X680), held in files encoded by the DER (Distinguished Encoding Rules).

The fundamental takeaway from that tech-dump is that X509 certificates are not in plain text, and there are multiple standards required in order to understand their contents.

So when someone says their certificate Distinguished Name is '/O=SomeUni/OU=SomeDept/L=group/CN=JohnSmith' ... that's not quite accurate. What they really mean is that there certificate DN is some set of objects that can be unambiguously matched to that ASCII text.

That happens because there are universally agreed mappings between the actually stored OID and the text representation of them (e.g. CN is OID 2.5.4.3).

Unfortunately, the agreement breaks down a bit for the emailAddress field; with some software mapping it to Email, and others to emailAddress. By the PKCS#9 standard, one could argue that it should be emailAddress - but that doesn't help us get software working.

Fortunatly, all of this is not a problem unless we want to store certificate DN's in ASCII, _and_ want to have email addresses in the DN.

Yeah, you can see where this is going, can't you?

In the UK, Jens has been working to allow us to not have them in DN's. However, in the short term, they are present.

One particular case where ASCII representations of the DN are used is in LSC files - which are used to authenticate VOMS servers. What happens is if the VOMS server DN matches the DN in the LSC file, and the cert was signed by the CA DN in the LSC file, _and_ the certificate chain is signed by a trusted root, then it's valid. This process means that we don't need to distribute lots of VOMS server certs, just the root CA's, and a small note (that shouldn't change over renewals) of the server DN.

I've been tidying up our ARC install here, and during the process managed to break things. Not unusual for me, (one of the reasons I avoid tiding at all costs!), but this one was quirky. I'd put the vomsdir under CFEngine control, so that it was sync'd with all the other servers, and suddenly it stopped accepting the scotgrid VO.

Root cause, as if you can't guess by now, LSC file, and the emailAddress. Looks like the gLite stack expects it one way, and ARC the other. Of course, by the time you read this, that's probably been fixed somewhere, but not in the version we had installed.

It turns out that there's one trick in LSC files that saves this case. Let me put the LSC file in here:

/C=UK/O=eScience/OU=Glasgow/L=Compserv/CN=svr029.gla.scotgrid.ac.uk/Email=grid-certificate@physics.gla.ac.uk
/C=UK/O=eScienceCA/OU=Authority/CN=UK e-Science CA
------ NEXT CHAIN ------
/C=UK/O=eScience/OU=Glasgow/L=Compserv/CN=svr029.gla.scotgrid.ac.uk/emailAddress=grid-certificate@physics.gla.ac.uk
/C=UK/O=eScienceCA/OU=Authority/CN=UK e-Science CA



The 'NEXT CHAIN' line lets one put multiple entries in the file. However, it appears that ARC isn't reading multiple, only the first one. So, in this case, I put the ARC friendly one first, so it matches fine - and the gLite stack tries again, finds the second, and thus suceeds.

Imporant notes: I can't find anyone else with a field report of NEXT CHAIN working in the gLite stack. This is such a field report. It doesn't appear to work with ARC.

Wednesday, December 21, 2011

Batch system juggling

We've been a bit quiet up here recently. This is normally a sign of either nothing interesting happening, or entirely too many interesting things happening. Opinions on that may divide, but I think it's closer to the latter...

One of the recent bits of fun that occurred was with our batch server. This story actually starts a long time ago; about this time last year. At that point, we started to get intermittent memory errors from the Torque server - corrected by ECC - but that's generally a sign that the RAM's about to fail. Given that the batch server is single point of failure for a site, that's not a good thing.

So I spent some time preparing a spare box, and being ready to move the batch system over, in case it failed over the winter break. Which, after all that prep, it didn't, and the errors stopped. On the expectation that the current hardware was nearing end of life, we ordered a new box early this year, and have had it sitting in a machine room for a while.

Unfortunately we didn't get time to have it running a tested batch system until our power supply started to ... well, insert colourful metaphor here, describing the 8 months where we were affected by lack of power.

Power got to stable supply in September, and so to catch up on things. One of the things we got around to was software versions. Whilst we didn't intent to update the Torque version, and managed to avoid it for a bit, the gLite developers eventually managed to sneak the update past us as part of an ordinary gLite update. Strictly, this didn't affect the batch server, just all the CE's, making them incompatible with the previous version of Torque.

Whilst a clever manoeuvre, reminiscent of Odysseus' Pony, it did leave us with a conundrum of either reverting the gLite update, or running forward with it. Neither were options of good character, but running forward did have some actual documentation; hence it was full speed ahead.

Which worked out well enough. The Torque 2.5.7 packages were set to use Munge, so getting that installed and tested as a first step helped it go smoothly. To preserve compatability in file locations, we used /etc/sysconfig/pbs_mom to put the pbs working directories in the same place as previously - meaning we didn't have to reconfigure any other tools.

What didn't go so smoothly was the memory leak in the server.

Which gave it a runtime of around 36 hours between crashes. Actually, not even crashes - we found that the pbs_server process hit either


12/05/2011 10:19:12;0080;PBS_Server;Req;req_reject;Reject reply code=15012(PBS_Server System error: No child processes MSG=could not unmunge credentials), aux=0, type=AlternateUserAuthentication, from tomcat@svr021.gla.scotgrid.ac.uk

or

10/29/2011 18:11:24;0001;PBS_Server;Svr;PBS_Server;LOG_ERROR::Cannot allocate memory (12) in send_job, fork failed


and then sat around moaning. Had it crashed hard, then the auto-restart would have caught it. Ho, hum, one for the Fast Fail philosophy there.


By this point, my proof reader is pointing out that I started off talking hardware, and now talking software. Punchline is that the new server that we never got a chance to use has a lot more RAM than the old server. Therefore we wanted to move the server from the old hardware to the new, to give it a lot more RAM space. That won't fix the memory leak, but will mitigate the problem a bit.

Conventionally, this would involve draining the cluster, repositioning the CE's and then starting up everything again. Had we done that, this blog post would be over now.

Instead, we did a rolling update. This let us move things over without having to do a full drain. The biggest problem with a full drain is that, while most of the jobs finish within a shorter period of time that then limit, there are always some that take the full duration. This leaves us with an empty cluster, doing nothing, for 24 hours or so, wainting on a couple of jobs to finish.

So, instead, by moving things in small batches, then we can keep most of the nodes working, and thus get more work out of things. Step zero is to disable cfengine, otherwise it tends to try and 'fix' things part way through.

Step one is to drain a CE, which we did over a weekend, and a small number of nodes, which we put offline on the Sunday morning.

Come Monday, I set up and tested basic operations with the new batch server, and then moved the freed up nodes across to it. Once those were tested (which shook out a couple of issues about versioning of some libs), point the CE at the new batch server, and then run a test job though it. (It turns out that Atlas are fast enough to sneak some pilots through a 2 minute window for a test job. However, only a few, so they actually functioned as effective tests, without compromising the site if they failed).

After that, it's time to offline another CE, and then some more nodes, and start moving nodes over when they were empty. In the end I scripted this:


#!/bin/sh

NODE=$1
RUNNING=$(qstat -n -1 | grep $NODE | wc --lines)

if [ "x${RUNNING}" != "x0" ]
then
echo $NODE: Still $RUNNING jobs going, skipping
exit 2
fi

CORES=$(qmgr -c "print node ${NODE}" | grep "np = " | cut -d= -f2)

FROM=svr666
TO=svr999

echo $NODE: Moving to ${TO} with ${CORES} cores

ssh ${TO} "~/addNode.sh ${NODE} ${CORES}"

ssh ${NODE} "service pbs_mom stop"
scp config.mom.svr666 ${NODE}:/var/spool/pbs/mom_priv/config
ssh ${NODE} "service pbs_mom start"

ssh ${FROM} "~/deleteNode.sh ${NODE}"


In theory one can run qmgr remotely, rather than ssh-ing to the batch servers and running a script. In practice, with the different versions of Torque, I couldn't get that to work. Note the automation of the mom config switch as well; and that this script checks that the node is empty.

This reduced the gradual move of nodes to a process of croning the script, and offlining nodes occasionally.

The net result was that we were operating at around 80% capacity for 48 hours, and it was all rather uneventful - in a good way. The final step was to update cfengine config and re-enable it.

One of the plus points of the above script is that it should be simple to adapt to two distinct batch systems; which means if we end up moving away from Torque, we should be able to do that without downtime too.

Friday, September 23, 2011

Leaving Lyon

The EGI Tech Forum is winding down, with only a few talks remaining. It's been a great meeting, with a wide range of talks on all areas of Grid Computing. Lots to think about and new ideas to try out!

Wednesday, September 21, 2011

Scotgrid goes South

Last we week attended the bi-annual GridPP Collaboration meeting.
The venue this time was CERN itself and the meeting was, as ever, incredibly useful.

We were lucky enough to have presentations from the Experiments, the LHC, EGI and the WLCG community as well as presentations from across the UK collaboration.

A full programme of the meeting is available here:

http://www.gridpp.ac.uk/gridpp27/



Above is a picture of our own Dr Crooks presenting on the Glasgow Security Model

Monday, September 19, 2011

EGI Tech Forum 2011

Bonjour Lyon!

After last week's GridPP 27 meeting in CERN, this week we are in Lyon for the 2011 EGI Tech Forum, running from Monday until Friday this week. You can follow the Forum online using some of the links here.

More later - time now to find some coffee before the first session...

Thursday, August 25, 2011

Busy Disks

After checking a test 10 gig Disk Server deployment we uncovered an interesting pattern in storage network activity and how our 10 Gig switch copes with multiply connections at 10 Gigabit. The captures below were taken over a 5 minute window of operation and show just how bursty the traffic patterns from these devices can be.

The graphs show all interfaces on our Dell 8024F and the measurement window is in Mbps. The order is top to bottom with the initial capture at the top.




While the Disk servers have been hammering away the round trip time intra room has been on average 0.40 msec between devices as the CPU on the core Dell seems more than happy to be handle these loads as its utilisation is approximately 20% presently.

We are planning to enable QOS metrics on disk server traffic shortly to test the response times on QOS and Non-QOS disk servers.