Showing posts with label scotgrid-gla. Show all posts
Showing posts with label scotgrid-gla. Show all posts

Friday, November 19, 2010

Second Cream CE for Glasgow First steps

We are currently in the process of installing a second Cream CE at Glasgow. This will replace one of the LCG CEs at Glasgow. As this is my first major service install since joining ScotGrid and the Gridpp project at the end of August I thought I would share the process for this type of service change with the wider community.


The first steps undertaken by myself was to drain the current LCG-CE to prepare it for the new install, the commands are shown below.



" For multiple CE's with shared queues. Edit the gip file on the CE you wish to drain. This blocks WMS submission: 

vim /opt/lcg/libexec/lcg-info-dynamic-pbs

change: push @output, "GlueCEStateStatus: $Status\n"
to: push @output, "GlueCEStateStatus: Draining\n" "


" on the batch machine: vim /etc/hosts.equiv comment out the machine you wish to stop accepting jobs and restart maui: 

svr016:~# cat /etc/hosts.equiv
svr021.gla.scotgrid.ac.uk
#svr026.gla.scotgrid.ac.uk "
 
However the GOCDB was not updated by myself to indicate scheduled 
downtime for this service change and after a GGUS ticket this was quickly
rectified.  We are waiting on the jobs to drain from the LCG-CE just now
 before continuing with the install early next week.

Wednesday, January 10, 2007

After updating to the latest gLite release (r11) on the old scotgrid site, the site started to fail JS with the much maligned message "Unspecified job manager error".

It turned out to be the classic WN cannot ssh back to CE to transfer job output. After a bit of poking around and playing I decided that for a site which has just 48 hours of run time left before it gets shut down it was just not worth the hassle of trying to debug this - so I have closed the queues and put the site into downtime, where it can languish until it is withdrawn on Friday.

Farewell, scotgrid-gla!

Tuesday, December 19, 2006

The old scotgrid production site (scotgrid-gla) has been failing on SAM rm tests for a few days because we have completely run out of storage! After the failure of the large RAID array we only had a single array with 1.6TB. It looks from the SRM logs that both ATLAS and ZEUS have been trying to add more files than we can cope with.

I have managed to secure a few 10s of GB by deleting some old dteam jobs. I have also reactivated a small pool on se2-gla itself and reserved it for ops.

I really need to do an EGEE broadcast today announcing the closure of the old site. Unfortunately the queue length is 200 hours - so we cannot close before Xmas.