Last we week attended the bi-annual GridPP Collaboration meeting.
The venue this time was CERN itself and the meeting was, as ever, incredibly useful.
We were lucky enough to have presentations from the Experiments, the LHC, EGI and the WLCG community as well as presentations from across the UK collaboration.
A full programme of the meeting is available here:
http://www.gridpp.ac.uk/gridpp27/
Above is a picture of our own Dr Crooks presenting on the Glasgow Security Model
Showing posts with label UKI-SCOTGRID-DURHAM. Show all posts
Showing posts with label UKI-SCOTGRID-DURHAM. Show all posts
Wednesday, September 21, 2011
Friday, May 08, 2009
ScotGrid Updates
Glasgow:
- Mike enabled pilot roles for both ATLAS and LHCb. He will also work on a parser which digests torque logs and gives the accounting figures in HEP-SPEC2006.
- Dug has been tracking down problems and discovering more about the LCG-CEs failure modes than he ever wanted to know (double job running from comms problems all down the line between ganga, wms, CE and batch system).
- Stuart has been optimising the cleanup of shared disk areas, which were cramping our style by sending the main nfs server into serious i/o wait for 20 hours in the day.
- Sam has installed a small test xrootd server - hopefully I will start running some analysis jobs against it soon to test it out.
- We reviewed our fairshares in advance of STEP09 to make sure each group was getting their due. We dropped most of our opportunistic VOs down to 1%.
- I discovered a jolly wheeze in Maui to use QOS to help bind the three different ATLAS fairshares into one QOS unit, with its own fairshare. This gives ATLAS sub-groups a fairshare advantage if the total ATLAS usage is under the total ATLAS target. Goes like this:
GROUPCFG[atlas] FSTARGET=10 MAXPROC=2000,2000 QDEF=atlasDurham:
GROUPCFG[atlasprd] FSTARGET=21 MAXPROC=2000,2000 QDEF=atlas
GROUPCFG[atlaspil] FSTARGET=11 MAXPROC=2000,2000 QDEF=atlas
QOSCFG[atlas] FSTARGET=42+
- Running well, but we decided not to implement the ATLAS pilot role (no intention to really support ATLAS analysis - they don't have the disk) and the LHCb pilot role is optional.
- Did the HEP-SPEC2006 benchmark on their nodes and got 67.82 for their Xeon L5430s (2.66GHz).
- To ward off less efficient user jobs we deleted ATLAS AOD - should see them only doing production for now.
- APEL publishing problem fixed.
- Steve plans to replace the ancient gLite 3.0 CE with a spiffy new gLite 3.1 one.
Labels:
ECDF,
UKI-SCOTGRID-DURHAM,
UKI-SCOTGRID-GLASGOW
Monday, April 20, 2009
Victim of our own success
I was wondering why Glasgow was not getting more activated ATLAS production jobs and eventually tracked it down to the fact that our cache disk area, ATLASPRODDISK, was almost full with only 170GB free space left. Panda was very sensibly not sending us more jobs until we had somewhere to put the outputs!
I quick whirl with dpm-updatespace and I increased PRODDISK from 2TB to 5TB, which should see us good.
I also discovered that Durham was missing some ATLAS releases, which was why they were missing out on ATLAS jobs today - installations now triggered.
I quick whirl with dpm-updatespace and I increased PRODDISK from 2TB to 5TB, which should see us good.
I also discovered that Durham was missing some ATLAS releases, which was why they were missing out on ATLAS jobs today - installations now triggered.
Labels:
ATLAS,
UKI-SCOTGRID-DURHAM,
UKI-SCOTGRID-GLASGOW
Monday, August 25, 2008
SAM Failures across scotgrid: Someone else's problem
All 3 scotgrid sites have just failed the atlas SAM SE tests (atlas_cr, atlas_cp, atlas_del) as have quite alot of the rest of the UKI-* sites.
Once again this isn't a Tier-2 issue but an upstream problem with the tests themselves
ATLAS specific test launched from monb003.cern.ch
Checking if a file can be copied and registered to svr018.gla.scotgrid.ac.uk
Once again this isn't a Tier-2 issue but an upstream problem with the tests themselves
ATLAS specific test launched from monb003.cern.ch
Checking if a file can be copied and registered to svr018.gla.scotgrid.ac.uk
------------------------- NEW ----------------
srm://svr018.gla.scotgrid.ac.uk/dpm/gla.scotgrid.ac.uk/home/atlas/
+ lcg-cr -v --vo atlas file:/home/samatlas/.same/SE/testFile.txt -l lfn:SE-lcg-cr-svr018.gla.scotgrid.ac.uk-1219649438 -d srm://svr018.gla.scotgrid.ac.uk/dpm/gla.scotgrid.ac.uk/home/atlas/SAM/SE-lcg-cr-svr018.gla.scotgrid.ac.uk-1219649438
Using grid catalog type: lfc
Using grid catalog : lfc0448.gridpp.rl.ac.uk
Using LFN : /grid/atlas/dq2/SAM/SE-lcg-cr-svr018.gla.scotgrid.ac.uk-1219649438
[BDII] sam-bdii.cern.ch:2170: Can't contact LDAP server
lcg_cr: Host is down
+ out_error=1
+ set +x
-------------------- Other endpoint same host -----------
Labels:
ATLAS,
monitoring,
nagios,
SAM,
UKI-SCOTGRID-DURHAM,
UKI-SCOTGRID-ECDF,
UKI-SCOTGRID-GLASGOW
Tuesday, June 24, 2008
Powercut at Durham
Ok, first blog post here...
On sunday, our machine room's UPS caused a brief power failure, which unfortunately tripped some breakers so we had to call out the electricians before we could start restoring service.
The UPS will take some time to repair, so Durham will be at risk for a while.
On the plus side, the changes involved with the SE rebuild have been proven to survive a reboot!
EDIT:
Just realised that I forgot to say that the site is back online, and has been since sunday evening, it is just currently at mercy of the power company.
On sunday, our machine room's UPS caused a brief power failure, which unfortunately tripped some breakers so we had to call out the electricians before we could start restoring service.
The UPS will take some time to repair, so Durham will be at risk for a while.
On the plus side, the changes involved with the SE rebuild have been proven to survive a reboot!
EDIT:
Just realised that I forgot to say that the site is back online, and has been since sunday evening, it is just currently at mercy of the power company.
Monday, June 23, 2008
Durham SE Issues
Durham suffered a complete SE failure last week. A RAID card failure took down the old SE, gallows, and then an LVM metadata corruption took out the new disk server on se01.
The list of lost ATLAS files has been reported (https://savannah.cern.ch/bugs/?38037) and we're waiting for the catalog to be cleaned up to restart production here (well, when there are any jobs to run).
We took the opportunity to retire gallows and now se01 is the sole SE at Durham. It should suffice for ATLAS production where we only need a few TB cache anyway.
In the meantime there was a power outage in the Durham machine room over the weekend. David had to get the university to reset some breakers but things seem to be running well now.
The list of lost ATLAS files has been reported (https://savannah.cern.ch/bugs/?38037) and we're waiting for the catalog to be cleaned up to restart production here (well, when there are any jobs to run).
We took the opportunity to retire gallows and now se01 is the sole SE at Durham. It should suffice for ATLAS production where we only need a few TB cache anyway.
In the meantime there was a power outage in the Durham machine room over the weekend. David had to get the university to reset some breakers but things seem to be running well now.
Friday, June 13, 2008
Problems up north
We have two major problems in ScotGrid right now:
ECDF: Have been failing SAM tests for over a week now. The symptom is that the SAM test is submitted successfully, runs correctly on the worker node, but then job outputs never seem to get back to the WMS, so eventually the job is timed out as a JS failure. As usual we cannot reproduce the problem with dteam or ATLAS jobs (in fact ATLAS condor jobs are running fine) so we are hugely puzzled. Launching a maual SAM test throught the CIC portal doesn't help because the test gets into the same state and hangs for 6 hours - so you cannot submit another one. Sam has asked for more network ports to be opened to have a larger globus port range, but the network people in Edinburgh seem to be really slow in doing this (and it seems it is not the root cause anyway).
Durham: Have suffered a serious pair of problems on their two SE hosts. The RAID filasystem on the headnode (gallows) was lost last week and all the data is gone. Then this week the large se01 disk server suffered an LVM problem and we can no longer mount grid home areas or access data on the SRM. Unfortunately Phil is on holiday, David is now off sick and I will be away on Monday - hopefully we can cobble something together to get the site running on Tuesday.
Thankfully, dear old Glasgow T2 is running like a charm right now (minor info publishing and WMS problems aside). In fact our SAM status for the last month is 100%, head to head with the T1! Fingers crossed we keep it up.
ECDF: Have been failing SAM tests for over a week now. The symptom is that the SAM test is submitted successfully, runs correctly on the worker node, but then job outputs never seem to get back to the WMS, so eventually the job is timed out as a JS failure. As usual we cannot reproduce the problem with dteam or ATLAS jobs (in fact ATLAS condor jobs are running fine) so we are hugely puzzled. Launching a maual SAM test throught the CIC portal doesn't help because the test gets into the same state and hangs for 6 hours - so you cannot submit another one. Sam has asked for more network ports to be opened to have a larger globus port range, but the network people in Edinburgh seem to be really slow in doing this (and it seems it is not the root cause anyway).
Durham: Have suffered a serious pair of problems on their two SE hosts. The RAID filasystem on the headnode (gallows) was lost last week and all the data is gone. Then this week the large se01 disk server suffered an LVM problem and we can no longer mount grid home areas or access data on the SRM. Unfortunately Phil is on holiday, David is now off sick and I will be away on Monday - hopefully we can cobble something together to get the site running on Tuesday.
Thankfully, dear old Glasgow T2 is running like a charm right now (minor info publishing and WMS problems aside). In fact our SAM status for the last month is 100%, head to head with the T1! Fingers crossed we keep it up.
Labels:
CE,
SAM,
Storage,
UKI-SCOTGRID-DURHAM,
UKI-SCOTGRID-ECDF
Friday, September 14, 2007
Durham News
Phil's procured a 15TB disk server for their DPM. This is now ready to go and we should get it configured and added to the Durham SRM next week.
Sunday, September 02, 2007
Durham Problems
Durham had a 2 day downtime last week as their network was down. Their central services were having a significant outage. A knock on effect from this seemed to be DNS issues which affected the site at the end of the week in a way that was not explained to me.
However, by the weekend the site was back up and running, so hopefully everything's ok now!
However, by the weekend the site was back up and running, so hopefully everything's ok now!
Friday, August 17, 2007
Durham Network Outage
Durham suffered a network outage last night, losing all connectivity to the outside world. The grid systems themselves recovered well this morning, but we suffered about 7 hours of downtime.
Tuesday, August 14, 2007
DPM Gridftp Resource Consumption

Durham was suffering from excessive resource consumption, from "hung" dpm.gsiftp connections from ATLAS transfers. Because of the way that gridftp v1 servers work huge buffers were being held in memory, leading to resource exhaustion on the machine and a subsequent crash.
Phil and I discussed this, and I noticed that the active network connections were to the RAL FTS server, not to the source SRM, so it looked like it was the control channel which was hung open, not the data channel.
Greig had a look on Glasgow's servers and discovered the same problem, but we were relatively unaffected due to the whopping 8GB of RAM we have in each disk server (and by having 9 disk servers, presumably). Cambridge also reported problems.
The issue is being looked at by the DPM developers, but for the moment Phil's had to write a cron script to kill off the hung ftps to keep gallow's head above water.
Labels:
ATLAS,
Data Management,
DPM,
UKI-SCOTGRID-DURHAM
Monday, July 23, 2007
Holiday Time: Durham Survives

Phil's been away for the last three weeks, with responsibility for Durham falling between myself (to respond to tickets and advise on grid problems) and Lydia (on the ground to press the buttons). This has worked pretty well - we have managed to deal with helmsley locking up (week 3 in the graph) and needing rebooted and a period of scheduled downtime (week 4, which revealed a problem in downtime synchronisation between SAM and the GOC).
A sterner test will come at Glasgow in a fortnight when Andrew and I are both on holiday and are more or less uncontactable. Time for icons and prayers?
Friday, July 20, 2007
SAM Gets Downtime Wrong?


Durham were in downtime from Wednesday -> Thursday, but SAM thinks they were in downtime from Thursday -> Friday. Doubtless this happened because I made the initial mistake (out by 1 day!), then edited the downtime to bring it forward. But SAM did not pick up the change.
I've raised a GGUS ticket - after all I'm always editing downtimes!
Monday, June 04, 2007
SAM tests changed VOMS Role (without warning!)
We, and a large fraction of the rest of the grid, started to fail replica management tests late on Friday night. At first I thought it must be a catalog problem at cern, so I raised a ticket. However, it turned out that what had actually happened was that the VOMS role used to submit the SAM tests had changed. This caused DPM to map the SAM tester DN into a different group - who then did not have permission to write into the default generated directory for lcg-cr.
This change was made completely unannounced, and I suspect without any real thought as to the implications for sites using DPM 1.6.3 and earlier.
Maarten Litmath helpfully posted a fix-up script on LCG-ROLLOUT, which uses ACLs to grant suitable privileges to lcgadmin and production roles for each supported VO, which I applied to Glasgow and Durham at about midnight last night (I was surely violating cardinal rules of sysadmining, but I couldn't see how it would cause harm - and this time I got away with it). This did fix the problem.
I'm really annoyed that this though. Changes like this should never, ever be made on a Friday! (It seemed the change actually came through at ~10am, but didn't break until midnight, when the next YYYY-MM-DD directory needed to be created.) In addition several people have commented that the fix is to upgrade to DPM 1.6.4 - despite the fact that this is broken in gLite 3.0r25 in two significant ways!
Grrrr. I just hope they don't ask us to explain these SFT failures - they shall have a piece of my mind... (I sound just like my Mum, when she was annoyed - see what the grid's doing to me!).
This change was made completely unannounced, and I suspect without any real thought as to the implications for sites using DPM 1.6.3 and earlier.
Maarten Litmath helpfully posted a fix-up script on LCG-ROLLOUT, which uses ACLs to grant suitable privileges to lcgadmin and production roles for each supported VO, which I applied to Glasgow and Durham at about midnight last night (I was surely violating cardinal rules of sysadmining, but I couldn't see how it would cause harm - and this time I got away with it). This did fix the problem.
I'm really annoyed that this though. Changes like this should never, ever be made on a Friday! (It seemed the change actually came through at ~10am, but didn't break until midnight, when the next YYYY-MM-DD directory needed to be created.) In addition several people have commented that the fix is to upgrade to DPM 1.6.4 - despite the fact that this is broken in gLite 3.0r25 in two significant ways!
Grrrr. I just hope they don't ask us to explain these SFT failures - they shall have a piece of my mind... (I sound just like my Mum, when she was annoyed - see what the grid's doing to me!).
Labels:
DPM,
SAM,
UKI-SCOTGRID-DURHAM,
UKI-SCOTGRID-GLASGOW,
voms
Monday, April 30, 2007
Durham Enable Fairshares
Mark, Phil and Nigel have enabled maui fairshares on the Durham cluster.
This is excellent news and should help them get better numbers on, e.g., Steve Lloyd's ATLAS tests.
I'm not quite sure how fair shares and pre-emption interact - I assume that they're somewhat independent.
This is excellent news and should help them get better numbers on, e.g., Steve Lloyd's ATLAS tests.
I'm not quite sure how fair shares and pre-emption interact - I assume that they're somewhat independent.
Wednesday, April 04, 2007
Defining SPEC Values for the Cluster
I had a long discussion with Mark about getting the SPEC values correct for Durham. There's no really good answer to this apart from go to the SPEC Website and try and find machines with the same processor types and vintage as your own (ideally with the same motherboard). N.B. One should really use the "base" values - these have a conservative set of compiler flags so are more appropriate for pre-compiled EGEE applications - the peak values enable all the bells and whistles on the compiler.
I was also prompted to look at the numbers I had put in for the new Glasgow cluster. Here we have Opteron 280s. There are now 10 measurements for the SI2K of these machines - these are all very close and average to 1533, so that's what I have now put (up slightly from 1450). The FP2K values have a bigger spread (different chipsets?), but in the absence of any guide I again took the average, which was 1770.
I also noticed that CPU2000 has now officially been retired - replaced by CPU2006. This is going to be a problem as CPU2000 will not be available for newer machines, but CPU2006 will not be available for older ones. How do you express that in your JDL?
I was also prompted to look at the numbers I had put in for the new Glasgow cluster. Here we have Opteron 280s. There are now 10 measurements for the SI2K of these machines - these are all very close and average to 1533, so that's what I have now put (up slightly from 1450). The FP2K values have a bigger spread (different chipsets?), but in the absence of any guide I again took the average, which was 1770.
I also noticed that CPU2000 has now officially been retired - replaced by CPU2006. This is going to be a problem as CPU2000 will not be available for newer machines, but CPU2006 will not be available for older ones. How do you express that in your JDL?
Labels:
Accounting,
SPEC,
UKI-SCOTGRID-DURHAM,
UKI-SCOTGRID-GLASGOW
Monday, March 26, 2007
Gatekeeper Troubles at Durham
Durham went on the blink at about 1am today - suddenly failing JL. The error message was the usual erudite globus effort: Got a job held event, reason: Globus error 3: an I/O operation failed.
Well, it looked straight forward enough - it's an i/o error, right. I found a lot of hints on google that this was caused errors in transferring in the sandbox. So check gridftp, home directory quotas, etc. Mark and I spent lots and lots of time on this, checking different things, becoming more and more confused (ok, so gridftp of a file works, can I make a directory using edg-gridftp-mkdir? have we restarted the gatekeeper properly? what's bound to ports 2811 and 2119? etc., etc.).
In the end we just could not fathom what had gone wrong, so I suggested to Mark that he email LCG-ROLLOUT and TB-SUPPORT.
Maarten Litmath pointed us to a GOC Wiki article which also said that this i/o error could occur when the CE was short of memory. I have found the culprit code in the l_check_memory function in Helper.pm - it produces a failure if the free memory (swap + physical) on the CE is less than 20% of the total. However, this error is not passed up the stack properly (in fact the code in queue_submit() returns undef) and so an entirely misleading error is passed back which wasted hours of our time. Grrrr.
I was reminded of a Alice in Wonderland...
'When I use a error message,' Humpty Dumpty said, in rather a scornful tone, `it means just what I choose it to mean -- neither more nor less.'
`The question is,' said Alice, `whether you can make error messages mean so many different things.'
`The question is,' said Humpty Dumpty, `which is to be master -- that's all.'
Alice was too much puzzled to say anything; so after a minute Humpty Dumpty began again. `They've a temper, some of them -- particularly Globus errors: they're the proudest - batch system errors you can do anything with, but not Globus errors - however, I can manage the whole lot of them! Impenetrability! That's what I say!'
`Would you tell me please,' said Alice, `what that means?'
I have submitted a GGUS ticket - these things won't improve unless they are complained about: https://savannah.cern.ch/bugs/index.php?25048.
Well, it looked straight forward enough - it's an i/o error, right. I found a lot of hints on google that this was caused errors in transferring in the sandbox. So check gridftp, home directory quotas, etc. Mark and I spent lots and lots of time on this, checking different things, becoming more and more confused (ok, so gridftp of a file works, can I make a directory using edg-gridftp-mkdir? have we restarted the gatekeeper properly? what's bound to ports 2811 and 2119? etc., etc.).
In the end we just could not fathom what had gone wrong, so I suggested to Mark that he email LCG-ROLLOUT and TB-SUPPORT.
Maarten Litmath pointed us to a GOC Wiki article which also said that this i/o error could occur when the CE was short of memory. I have found the culprit code in the l_check_memory function in Helper.pm - it produces a failure if the free memory (swap + physical) on the CE is less than 20% of the total. However, this error is not passed up the stack properly (in fact the code in queue_submit() returns undef) and so an entirely misleading error is passed back which wasted hours of our time. Grrrr.
I was reminded of a Alice in Wonderland...
'When I use a error message,' Humpty Dumpty said, in rather a scornful tone, `it means just what I choose it to mean -- neither more nor less.'
`The question is,' said Alice, `whether you can make error messages mean so many different things.'
`The question is,' said Humpty Dumpty, `which is to be master -- that's all.'
Alice was too much puzzled to say anything; so after a minute Humpty Dumpty began again. `They've a temper, some of them -- particularly Globus errors: they're the proudest - batch system errors you can do anything with, but not Globus errors - however, I can manage the whole lot of them! Impenetrability! That's what I say!'
`Would you tell me please,' said Alice, `what that means?'
I have submitted a GGUS ticket - these things won't improve unless they are complained about: https://savannah.cern.ch/bugs/index.php?25048.
Monday, March 19, 2007
New VOs at Glasgow: camont, gridpp, totalep
I've enabled the camont, gridpp and totalep VOs on the Glasgow cluster.
A complete description of how to do this is on the wiki.
This stupidest task must have been to write a script to reverse engineer a YAIM users.conf file from our group/passwd files so that the YAIM utility functions like users_getvogroup work. There's surely an easier way of doing that?
This afternoon I'll redo the RB and try and run some jobs through as a gridpp VO member.
A complete description of how to do this is on the wiki.
This stupidest task must have been to write a script to reverse engineer a YAIM users.conf file from our group/passwd files so that the YAIM utility functions like users_getvogroup work. There's surely an easier way of doing that?
This afternoon I'll redo the RB and try and run some jobs through as a gridpp VO member.
Labels:
camont,
GridPP,
totalep,
UKI-SCOTGRID-DURHAM,
VO
Monday, February 26, 2007
Durham Queue reogranisation
Last week Durham implemented new queue structure, we now have three grid queues -
- ops (for ops jobs)
- pheno (for the phenogrid VO)
- grid (for all other grid jobs)
Thursday, February 15, 2007
Steve Lloyd was still having problems writing into Durham's DPM. I nabbed an ATLAS user and got them to help me, which has revealed an incredibly weird problem:
If the user is unknown to the DPM then lcg-cr will fail, with the cryptic error "transport endpoint not connected". Attempting to srmcp reveals an authentication failure: "SRMClientV1 : CGSI-gSOAP: Could not find mapping for: USERS_DN". Somehow the srm daemon is refusing to authenticate the user properly. It's actually quite a deep problem in the GSI chain, because the daemon doesn't even get as far as logging anything about the connection.
Drilling further down, to globus-url-copy, then revealed a bizarre work around: a globus-url-copy command will create the mapping from the user's DN to a pool account. After this is done everything starts to work properly.
We've checked to see if it's a problem with VOMS proxies and it isn't. It also doesn't seem to affect DPNS itself - even users who can't copy in files can do a dpns-mkdir and they get a new entry in Cns_userinfo fine.
Very mysterious.
If the user is unknown to the DPM then lcg-cr will fail, with the cryptic error "transport endpoint not connected". Attempting to srmcp reveals an authentication failure: "SRMClientV1 : CGSI-gSOAP: Could not find mapping for: USERS_DN". Somehow the srm daemon is refusing to authenticate the user properly. It's actually quite a deep problem in the GSI chain, because the daemon doesn't even get as far as logging anything about the connection.
Drilling further down, to globus-url-copy, then revealed a bizarre work around: a globus-url-copy command will create the mapping from the user's DN to a pool account. After this is done everything starts to work properly.
We've checked to see if it's a problem with VOMS proxies and it isn't. It also doesn't seem to affect DPNS itself - even users who can't copy in files can do a dpns-mkdir and they get a new entry in Cns_userinfo fine.
Very mysterious.
Subscribe to:
Posts (Atom)