Would You Like More WLM Information In DB2 Accounting Trace – And How Would You Use It?

(Originally posted 2012-02-06.)

I was lucky enough to be in Silicon Valley Lab for DB2 BootCamp last week. There I ran into a DB2 developer I’ve worked very successfully with in the past – John Tobler.

(He’s the guy I look to for questions and issues with DB2 SMF data.)

We had a good discussion about something I’d personally like to see in DB2 Accounting Trace – more WLM information – and this post is as a result of this conversation.

Two salient pieces of information:

  1. Accounting Trace already has a field for WLM Service Class (QWACWLME) but it’s only filled in for DDF work.
  2. As Willie Favero pointed out in APAR Friday: WLM information is now part of the DISPLAY THREAD command the command now has some WLM information in it.

Putting these two together you come to the conclusion it might technically be possible to get more WLM information into Accounting Trace. That, of course, doesn’t mean it’s going to happen. I have to stress that before going any further. But it’s worthwhile thinking about what’s needed and how useful that would be to customers.

What Should Be Added?

Uncontroversially, I think, QWACWLME should be filled in with Service Class for all work types. I say "uncontroversially" because – if it can be done cheaply – it’s just using space that’s already in the record. I don’t know if it can be done cheaply, though.

More controversially because, taken together, they represent 18 additional bytes in each 101 record are:

  • WLM Workload
  • WLM Report Class
  • WLM Service Class Period

I think I could live without Workload but it seems a shame to exclude it.

As Willie points out Performance Index (PI) is also in the DISPLAY THREAD command but I think we can get that from RMF Workload Activity Report (SMF 72) and that’s probably a better place to get it from.

But the key question is “how useful and important would this extra information be to you?”

Let me outline three areas of use I can immediately see…

Understanding Not Accounted For Time

This time bucket is what you get when you subtract all the time buckets we know about from the headline response time. The two most important causes for this are CPU Queuing and Paging Delay.

If we calculate this time for a record and we know the (behaviour of) the WLM Service Class it’s in we can understand this time better. A bugbear of doing DB2 performance is just this: understanding whether work is subject to queuing or not. (For Paging Delay as a cause of Not Accounted For Time we could do much the same thing.)

Understanding The WLM Aspects Of DB2 Work

It would be useful to be able to break down the work coming into a DB2 subsystem by Service Class, Goal and Importance, wouldn’t it? In particular it would be nice to see the hierarchy of goals and importances, and to be able to relate the works’ WLM attributes to those of address spaces such as DIST and DBM1. (In the former case discovering that the TCB’s in the DIST address space were subject to pre-emption by the DDF work would be a blow.)

Correlating Service Class And Report Class For DDF Work

For non-enclave work I use the Report Class and Service Class in Type 30 to establish how these relate to each other (and what kind of work has which RC and which SC). I can’t do it for DDF work because there’s no usable Type 30 (i.e. with this kind of information in). If the 101 record had these both in you could extend the method.

(In case you wonder what I’m talking about see What’s In A Name?.)

This still doesn’t help us in the non-DDF enclave cases, of course.

Over To You

What do you think? I’ve listed three categories of value that immediately spring to mind (and that’s with the disbenefit of jetlag so maybe not that articulately expressed). But I’d really like to know if this would be of value to you – and to modify the proposal if you think you’d like something slightly different.

There’s no guarantee this will get done – and it’s a bit of an attempt at a “Social Requirements Gathering” process. But it’s worth debating in public, I think.

Haven’t We Been Here Before?

(Originally posted 2012-01-28.)

Well, some of us have. πŸ™‚

Well before we announced zEnterprise I thought it would be rolled out and adopted in a similar manner to Parallel Sysplex (and to many other technologies – whether mainframe or otherwise).

Reading zEnterprise Use Cases Start Rolling In I still think I’m right. And I will admit I needed to see something encouraging like this.

Back in the mid 1990’s we introduced Parallel Sysplex. In fact we started with Sysplex and then added the "Parallel" elements to it.

Adoption of Parallel Sysplex took a while. And hence the folklore and confidence in the value proposition took a while to take root.

If I were to list the things that needed working on to make Parallel Sysplex mainstream you might mistakenly think the same list (or even a similar sized list) of "to do’s" applied to zEnterprise. You can’t draw that conclusion. You can draw the "appropriately speedy train coming" parallel but that’s all.

But let’s revisit (a subset of) that list:

  • Performance and efficiency improvements.
  • More exploiters
  • More function
  • Enhanced Availability
  • Extra Instrumentation
  • Field – whether IBMer or customer or consultant or third-party vendor – experience

As I said, don’t take that list as a template for the way zEnterprise is going to evolve. But if you "squint" at the list some familiar themes emerge.

And the referenced blog post addresses one of these: Customer experience. Though I don’t manage the agendae for conferences it wouldn’t surprise me if we saw some "customer experience" presentations soon.

As a young Systems Engineer in the late 1980’s I saw a number of considerably simpler product function introductions. As those of us who were around all know there was a hurry on – at least from IBM’s perspective: Our competitive differentiator (and new product vs old differentiator) was new function we hoped customers would adopt quickly and really value. You can think of Hiperbatch if you like. But if you do I’d prefer you to think of the MVPG instruction (the hardware function it relied on) which was used by a number of other functions to cut CPU. I’m thinking primarily of VSAM LSR Hiperspace buffers here. And, while we’re at it how about ADMF? Both MVPG and ADMF were used together by DB2 Hiperpools – again to cut CPU.

The reason for detailing MVPG and ADMF is they had clear advantages for many customers – and still they took in excess of 18 months from announcement to widespread adoption. I’d say they were simple to implement as well.

I don’t think anyone would claim Parallel Sysplex or zEnterprise full functionality are quick or simple to implement: If you’re looking at the sheer sweep of what we’re doing I think that’s appropriate.

So, I think we’re in good shape: We’re now seeing implementations and I’m sure we’re going to see many more. And I do think the Parallel Sysplex analogy is a good one – in terms of choreography of adoption.

Sometimes I think those of us have been around have only the “we’ve been here before” perspective to offer. Actually I think we do have that. But, of course, I think we have a lot else besides to offer: Thinking about Systems and value as well as the “calmness” πŸ™‚ of knowing “this is how it goes”.

This is going to be fun – and fun soon. πŸ™‚

A Better Calibre of Kindling

(Originally posted 2012-01-23.)

You might consider it showing off if I mention I got a Kindle for Xmas. Feel free to. πŸ™‚ But I’d like to share my experience with you – as you might find it useful anyway.

First, I really like the Kindle as it stands. Mine is a Keyboard 3G one. I felt both the “keyboard” and 3G elements were important:

  • I surmised (correctly) I’d want to take notes.
  • I surmised (equally correctly) I’d want to be able to do things wherever I was that would need access to “Kindle Central”. (Actually, access at 35,000 feet will have to wait.)

I’ve found the basic act of reading on the Kindle to be at least as rewarding as reading paper books. I also appreciate putting an end to being engulfed by the rising tide of new books.

(In the house I seem to be the one that wants to keep books once I’ve read them. I’m also the one who doesn’t feel the need to complete a book if I’ve read it. So I have several books on the go at the same time on the Kindle and it’s kept track of where I am with them all. Yes, I know it’s called a bookmark so no distinct advantage there.)

I also appreciate the social aspect:

  • Sharing snippets via the Kindle website and posting links to them on Twitter. Some of you will have seen that – probably most of you given I propagate tweets to Facebook and LinkedIn.
  • I’m re-reading Terry Pratchett’s “The Colour Of Magic” and it’s nice to see “you and 5 people” against key quotes. I don’t know who these people are but already I feel kinship with them. πŸ™‚

Book delivery is pretty swift – which is much more than can be said of ordering paper books. And I’ve used the “try a sample” capability several times: With both positive and negative buying outcomes. I’m using the Amazon “Wish List” as my queue for acquiring books so I don’t necessarily buy immediately.

Calibre

There isn’t much need for curation but my tool of choice for doing so is Calibre which is available for Windows, Linux and OS X. (I run it on Linux and OS X (though others in the house have Windows and there’s one other Kindle in the house). It’s free and it’s very good. One tip: If you’re using it on Linux it’s probably best to install it directly, rather than going through e.g. Debian repositories. I say this because it’s frequently updated and the repositories seem to be way behind.

I used Calibre with my old Sony PRS-700 eBook reader – which I found to be unusably slow and hard to read. (The Kindle is neither of these.)

Calibre does a number of things for me. Most notably it lets me:

  • Convert books from other formats e.g. EPUB.
  • Download RSS / Atom “news” feeds and convert them to MOBI so I can read them on Kindle.
  • Edit metadata for books – such as titles and authors. (Mainly this is worthwhile for books that weren’t from the Kindle Store – as some of them have dubious spellings etc.)
  • (I actually don’t feel the need to have Calibre back up my Kindle – though it will do that as well)

Calibre has a lot of sophistication built into its conversion. I’ve yet to fully explore what it can do, for instance, to tidy up conversion of PDF documents. Page footers, for one, need removing on conversion.

One other thing: You can use Calibre in Batch Mode. That might well help with automation.

Project Gutenberg

I’ve known for a long time about Project Gutenberg. To quote from their website:

“Project Gutenberg offers over 38,000 free ebooks: choose among free epub books, free kindle books, download them or read them online.”

Two good things to note:

  • Project Gutenberg has a rigorous copyright checking process – so everything is out of copyright or otherwise in the public domain. I’m against ripping off authors, so this is a good thing.
  • The books are well formatted: eBook quality can vary enormously, to the point where books can be frustratingly hard to read (in the worst case).

Without listing the catalog I’d say you can find many classics there. The “usual suspects” like Chaucer, Shakespeare and Oscar Wilde are represented (all of which I have on my Kindle), along with many others. (I wish Raymond Chandler were there but the absence of his works probably means they’re still under copyright protection.)

Distributed Proofreaders

So, where do Project Gutenberg books come from? I can’t say this is true of all of them but many come from Distributed Proofreaders. The idea of this is that people sign up to proofread OCR’ed pages – one page at a time. I signed up to do this and worked on the first proofreading of two books. I’d never heard of the books before and the actual process was good as I found the books interesting in their own right.

The OCR process was pretty accurate but the proofreading was absolutely necessary. I think it might be possible to codify many of the errors in the OCR process as they were repeated.

There are several rounds of proofreading and so the results – books in Project Gutenberg – are very good. There’s a lot of emphasis on not correcting the spelling or punctuation, and on not editorialising.

More volunteers are needed. As I say I’ve enjoyed doing it.

Hacking

If you connect a Kindle to a PC or Mac (and I’ve done both) the Kindle shows up as a removable drive. The most useful thing you can do with it is to extract the ‘My Clippings.txt’ file. This contains all your bookmarks and annotations. It’s reasonably hackable: While it’s not XML (and I really wish it were) it has a simple-to-understand and easy-to-parse format in plain text.

One challenge I’d like to see someone meet is processing this file and creating Evernote notes. True you can get at your annotations etc from Amazon but I think there’s value in easing getting marked up passages into Evernote. Indeed I’d be pleased if Amazon and Evernote worked together to provide a slick “clip to Evernote” function for Kindles other than the Fire.

I have other hacking challenges else I’d work on this one – processing the file – myself. I know that doing it for Windows (and Linux under Wine) and for OS X would mean two separate pieces of code.

So Why Am I Still Carrying Around Paper Books?

It turns out I still have a few books to get through in paper format before I go “all electronic”. I also expect there to be incidences where someone gives me a book. I consider those to be “beyond my control”. πŸ™‚

One final thing: For another view (although a corroborative one) see Susan Visser’s blog posts on the subject.

Rough And Ready?

(Originally posted 2012-01-20.)

A couple of items from the world of music caught my attention recently – and there’s some commonality between them:

  • According to Dave Grohl of Foo Fighters’ blog post: Hey everybody, Dave here’

    “From day one, the idea for this record was to make something completely simple and honest, to capture that thing that happens when you put the 5 of us in a small room. No big production, just real rock and roll music: That’s why we decided to do it in my garage. We wanted to retain that human element, keep all of those beautiful imperfections: That’s why we went completely analog.”

and

Of course I have both the Foo Fighters album (Wasting Light) and Beyond Magnetic. I thoroughly enjoy them and their roughness in no way detracts from the value I get from them. In fact both these comments surprised me.

Now granted neither Foo Fighters nor Metallica are known for their subtlety. πŸ™‚ But they are known for being amongst the best bands active today.

There is of course another band of exceedingly high effectiveness: Queen. Now they are known for their subtlety (mostly). πŸ™‚ But they’ve not been all that active for many years – for obvious reasons. 😦

It turns out there’s quite a lot of stuff in the Queen vaults that never officially saw the light of day. The suggestion is it’s unfinished and therefore not to be released. I, like many other fans, have heard some of this. We tend to think most of it meets our releasability criteria. Take for instance a song called I Guess We’re Falling Out. If you listen to it it’s clearly unfinished but absolutely exquisite. Now whether it should be released finished or unfinished is a good question. But I think it should certainly see the light of day.

Now this post isn’t just a rail against Queen Productions. It is that πŸ™‚ but it’s also about the wider point:

When is something good enough to see the light of day?

I’m obviously not advocating shoddy work – and none of these three examples from music represent that. But sometimes throwing something Rough And Ready (the title of this post, complete with pun) out there is the right way to go. And sometimes it’s not.

  • When I put together new analysis code it’s prototypical. And it’s the commitment to refine it in the light of experience that’s key here. As is the appropriate level of tentativeness involved.
  • When I’m doing something where quality is critical it’s a different matter entirely.

This post isn’t profoundly philosophical πŸ™‚ but it’s an area I did some thinking about over the holiday season. This time no new code of any value emerged from the holiday. But this and a couple of other lines of thinking did. Maybe I’ll post about those soon.

Java’s Not The Only JVM-Based Language

(Originally posted 2011-12-18.)

JVM-based languages have an interesting property for z/OS programmers: They are zAAP-eligible.

As we all know zAAP Eligibility brings a number of benefits – including the licence charge benefits and the ability to run on a full-speed processor even when your general-purpose processors are subcapacity ones. (I’ll briefly mention zAAP-on-zIIP here for completeness, and then move on.)

You probably also know that recent zSeries and System z processor generations have very significantly boosted JVM performance. That’s a combination of JVM design improvements, processor speedups and JVM-friendly processor instructions. These are properties of the JVM and processor rather than the language or the javac compiler.

I’ve carefully avoided saying "Java" so far in this post, apart from obliquely in the previous sentence. That’s because this post explores the notion that anything that runs in the JVM can take advantage of all the above. Equally the usual considerations come into play – most notably native code (JNI) affecting eligibility and the startup cost for the JVM.

So what is a JVM? The acronym stands for "Java Virtual Machine". In reality it’s a bytecode interpreter – pure and simple. There’s nothing that says those bytecodes have to be created using the javac java compiler. Indeed there are a number of languages that create bytecode for the JVM.

Further, there are languages that are interpreted by java code – and hence also run in the JVM. My expectation is this would be slower than those that create bytecode. These languages include Javascript (via Rhino,as I mentioned here), Python via Jython, and Ruby via JRuby (which I personally haven’t explored yet).

And then there’s NetRexx. Which we’ll come to in a minute.

So, why the fascination with other JVM-friendly languages? First, when people talk about Modernisation on the mainframe there’s often a strong component of java in it. My take on Modernisation contains two elements I want to get across:

  • Java isn’t the only modern language. Indeed I’d hazard it wasn’t particularly modern. For fans of programming languages take a look at the languages I’ve already listed in this post. And this matters because people with enthusiasm and programming skill will often be conversant with these languages. Furthermore, lots of stuff I’d like to see run on the mainframe under z/OS is already available, written in these languages
  • As I, perhaps grumpily, state in discussions on e.g. Batch Modernisation, the point is to "kick the ball forward", whether that means java or not.

Β So, back to NetRexx. It’s not the only flavour of REXX available under z/OS Unix System Services. That much is well known. But it does run in the JVM – by compiling NetRexx programs to java source. This is different from the "bytecode" and "interpreted by a java" program approaches. The result is a java class or jar file, just as if you’d written it in java in the first place.

I uploaded the two necessary NetRexx jar files – from the distribution downloadable from here. These are NetRexxC.jar – the compiler – and NetRexxR.jar – the runtime. (I suspect you only really need NetRexxC.jar.) When you compile a NetRexx program you place NetRexxC.jar in you classpath and invoke java program org.netrexx.process.NetRexxC.

I wrote a simple NetRexx program – which uses (automatically imported) java classes: java.util.regex.Pattern and java.util.regex.Matcher. This program takes from the command line a search string, a replacement string, and a string to search-and-replace in. When I say "simple" the NetRexx program turns out to be much simpler, shorter and more understandable than the java equivalent. Here it is:

parse arg lookup replacement s
say Pattern.compile(lookup).matcher(s).replaceAll(replacement)

And that really is all there is to it. The "parse arg" and "say" instructions should look familiar to anyone who knows REXX. The rest is just stacked invocations of java classes.

As I read through the NetRexx language definition I could see a lot of advantages over traditional REXX. Two you’ve already seen – interoperability with java classes and running in the JVM (though the latter isn’t really a language definition benefit). Others included object orientation, more sophisticated switch statements, and "–" to start a comment on a line. So I think this is a better REXX. I noted only one incompatibility (and it might not even be incompatible): "loop" instead of "do" to start a loop.

Because Classic REXX can’t do Regular Expressions (though "parse" is nice) I experimented with invoking my NetRexx program (above) from classic REXX. I used BPXWUNIX (mentioned before here (and I’ve now corrected that post which incorrectly mentioned BPXWDYN instead) and in Hackday 9 – REXX, Java and XML ). Because the program uses the JVM I made sure to pass environment statements to BPXWUNIX setting up PATH for the JVM. This worked very well.

I could’ve used BPXWUNIX to call sed instead – for this use case. That probably would’ve been cheaper but I was proving I could call a NetRexx program from TSO REXX (in batch), passing in parameters and retrieving the result. Talking of "cheaper" I think it’s important to try and avoid transitions across the BPXWUNIX boundary: It’ll have a (non-zAAP) CPU cost and, if you’re using the JVM as NetRexx does, it’ll cost to set up the JVM and tear it down again afterwards. A pair of transitions with meaty application processing in between is going to be the most efficient.Β 

(The previous paragraph was conjectural: It would be nice to run some benchmarks on this one day. Anyone?)

So, I’m impressed (as you can tell) with NetRexx. I think it’s worth taking a look at – as indeed are the other JVM-based language implementations I mentioned. The point of this post is to demonstrate (yet again) there are choices – and considerations to go with them.

What’s In A Name?

(Originally posted 2011-12-16.)

This is the post I was going to write before the discussion that led to CICS VSAM Buffering arose. It’s about getting more insight into how WLM is set up and performing than RMF Workload Activity Report data alone allows.

I recognise some of this can be done with the WLM policy in hand. But this is about an SMF-based approach. (The piece you can’t do with SMF is discerning the WLM classification rules.) And the policy can’t answer questions about how systems actually behave.

There are two distinct problems I’ve worked on solving (relatively) recently. I share the outline of my solution to each of these with you here.

  • In RMF you can’t tell how Report Classes and Service Classes relate to each other: In some cases Report Classes break down Service Class data – often to the address space level. In some cases Report Classes coalesce information from multiple Service Classes. But you can’t see this linkage in RMF.
  • In RMF you can’t necessarily tell what runs in each Service Class. I say "necessarily" because you can tell some things about the nature of the work in a Service Class.

The "What’s In A Name?" in the title refers to the fact a Workload, Service Class or Report Class name is just a string of characters: Rhetorically it might be a "promise" but it’s not a mechanistic guarantee. So – to me at least – it’s worth knowing rather more.

Report And Service Class Relationships

SMF 72 Subtype 3 RMF Workload Activity Report data describes how Service Class Periods and Report Classes perform.

Type 30 Interval records (Subtypes 2 and 3) describe how address spaces perform.(Actually so do Subtypes 4 and 5, which are step-end and job-end records.) These records contain, amongst other things, WLM Workload, Service Class and Report Class names – for the address space. You can therefore use Type 30 to relate Workload and Service Class to Report Class. My code’s done this for some time.

Type 30 does not apply to Service Classes that don’t own address spaces. Two examples of this are DDF Transaction Service Classes and CICS Transaction Service Classes.

A related topic is which Service Classes are serving other Service Classes. For example CICS Region Services Classes and transaction Service Classes. Now this you can readily discern from SMF 72 alone. (And of course my code does that.)

What Work In A Service Class Is

(This piece relates equally to Report Classes.)

As I said, you can’t tell much about what a WLM Service Class covers from Type 72. So, as well as the correlation described above, my code uses Type 30 to flesh out what a Service Class is for. The key to this is the Program Name. For example CICS regions have PGM=DFHSIP. So a Service Class with just PGM=DFHSIPΒ  address spaces is just a CICS Region Service Class. Simple enough. Some are more complicated than others – perhaps necessitating the 16-character program name field which, for Unix, includes the last portion of the Unix program name.

You can play other games, too: The job name for a DB2 address space can be decoded to glean the subsystem it belongs to. Certain System address spaces have mnemonic Procedure names. And so on.Β 

From SMF 72 you can obtain the number of address spaces for a Service Class – 0 suggesting the Service Class doesn’t own any (see above). 1 suggests this class (possibly a Report Class) is there to provide more granularity. You can also get the number of address spaces In and the number Out-And-Ready. This can help you form a picture of e.g. "low use" address spaces in the Service Class.

This post is about sharing some of my experience of trying to extend the value that can be got out of SMF – beyond the obvious. Some of this will probably appear in my I Know What You Did Last Summer presentation – which I’m still hoping to complete soon. This also, by the way, explains why I’m so keen to get Type 30 data from you when you’re sending me RMF data. There really is a huge amount of value to be had.

CICS VSAM Buffering

(Originally posted 2011-12-16.)

Four score and seven years ago (or so it seems) πŸ™‚ the Washington Systems Center published a set of mainframe Data-In-Memory studies. These were conducted by performance teams in various IBM labs and were quite instructive and inspiring. I wish I could find the form number (and a fortiori a PDF version) for this book. Anyone? Even hardcopy would be really nice.

The reason I mention this is because of a thread in the CICS-L newsgroup overnight about the CPU impact of increasing the size of VSAM LSR buffers in CICS. I seem to recall that CICS / VSAM was one of the benchmarks written up in this orange book. The original poster wanted to know what the CPU impact profile was of increasing VSAM buffers. I think the study showed that there could be some CPU saving with bigger buffer pools. (Compare this with VIO in (then) expanded storage – which showed a net CPU increase for the technique.)

There are a number of points I would have raised in CICS-L but I’ll write them here instead – as most of you probably don’t read CICS-L:

  • I would not build a Data-In-Memory (DIM) case on CPU savings (though I would want to satisfy myself there wasn’t a significant net cost). I would build it on throughput enablement and response time decreases. This is true of any DIM technique.
  • The thread in CICS-L correctly identifies the need to be able to provision real memory to back the increase in virtual.
  • VSAM LSR buffers are allocated from within the virtual memory of the CICS address space. For most customers this isn’t an issue as the buffers are usually within 31-bit memory. (There is no 64-bit VSAM buffering.) But it’s still worth keeping an eye on CICS virtual storage (whether 31- or 24-bit) – perhaps using what’s in CICS SMF 110 Statistics Trace.
  • Back in the late 1980’s there was a tool – VLBPAA – that would analyse User F61 GTF Trace to establish the benefit of bigger buffer pools – at least in raw I/O reduction terms. The trace is still available and you could process it with DFSORT but it would be harder to predict buffering outcomes without VLBPAA. In fact I mention this in Memories of Batch LSR
  • One of the comments talked about hit ratios but I prefer to think of miss rates – or better still misses per transaction.

In general I find CICS VSAM LSR buffering insufficiently aggressive: As memory is generally plentiful these days (at least relative to the CICS VSAM LSR pool sizes I encounter) I think it’s appropriate for installations to consider big increases (subject to the provisoes above). Think in terms of doubling rather than adding 20%. And no 10MB total buffering is not aggressive. πŸ™‚

DB2 Accounting Trace And Unicode

(Originally posted 2011-12-12.)

As I said in this post I recently came across the need to handle Unicode when processing DB2 Accounting Trace (SMF 101). I was astonished not to have run into it before in all my many sets of customer data. So I had two things to do:

  • Understand the circumstances under which it happens – which isn’t just "be on Version 8 and it will happen automatically."

and

  • Figure out how to handle it when I see it. (i.e. when QWHSFLAG has the value x’80’ as I mentioned in the other post).

As you’d expect, I asked the customer what they had done to cause the generation of 101 records containing Unicode fields. The answer is that they’ve set parameter UIFCIDS in DSNZPARM to "YES". It turns out my friend Willie Favero had mentioned it in this blog post some time ago. Because the DB2 Catalog has Unicode in it in Version 8 it actually takes cycles to create 101 records without Unicode in: All the fields marked "%U" in the mappings in SDSNMACS have to be translated from Unicode to EBCDIC. If you code "UIFCIDS=YES" you avoid the cost of the translation.

But there’s an obvious downside: Any reporting against those fields (the ones marked "%U") needs to take that into account. But if you never (or rarely) look at e.g. the Package-level stuff you might prefer to write it in Unicode (or suppress IFCID 239 entirely). It’s probably a net saving in CPU, albeit a small one.

Which leads on to the second part of this: How did I handle the translation into something readable (EBCDIC being my primary encoding, at least on z/OS)? My interim take is to fix up my reporting REXX to check QWHSFLAGS and do the right thing. You can readily do that with the built-in TRANSLATE instruction. There is code knocking around on the Internet for the purpose. That got me through this study and I have reusable code I can use wherever I need to do the translation.

But this is not the only way, and perhaps not the best: My code reformats records (for historical reasons, mainly) in an assembler exit – as part of my database build process. It’s entirely feasible to do the translation there and then my database has everything readable. To do that you can use the TR (Translate) instruction. Of course you can use the same translation table (albeit with different syntax – one or so edit away) as in the REXX. But that’s a whole load more effort and potential fragility. I think I’ll defer that.

One other thing: Unicode isn’t necessarily 1-byte characters. But the sample I have is. Neither REXX TRANSLATE BIF nor TR instruction will handle multi-byte characters. So I could eventually come unstuck. And, no, I don’t know if these fields will contain multi-byte characters any time soon. Anyone in a position to comment?

| Fixed a glitch in the bulleted list at the top.Β 

CICS and Batch

(Originally posted 2011-12-09.)

In my experience there are two kinds of CICS installations: Those that take CICS down at night – to run the Batch – and those that don’t.

There is a loose correlation between what the data manager is and which approach is taken: VSAM-based CICS applications tend to be less 24×7 than DB2 ones, though it’s not that clear cut.

This post is about how you (really I) might glean how you run CICS vis-a-vis Batch, using SMF. Even if you know the principle of how you manage CICS regions the numbers should still be useful and not too onerous.

If you’re the kind of installation that takes CICS down the SMF 30 Job-End records (Subtype 4) will tell you when CICS started and stopped. (If you want to know when CICS came down in an unscheduled manner the same data applies.)

If you don’t expect CICS to come down often the SMF 30 Interval Records (Subtypes 2 and 3) will confirm the region is still up.

(The (ancient) SMF processor I use – SLR – has an Availability Reporting capability. I would expect and hope other tools have something similar. In any case it’s just another take on regular performance data. I’m considering playing with SLR Availability Reporting.)

For the purposes of this discussion I consider any address space running program "DFHSIP" to be a CICS region. There may be other program names of relevance, too.

My preferred means of display is a Gantt chart – my (also ancient) formatting tool – Bookmaster – doing very nicely in that regard. I got into Gantt charts for Batch Suite display but the technique is fine for online regions. I’m beginning to annotate Gantt charts with commentary – a technique you may find useful. (I might post soon on how I’m doing the annotation – as that has been an interesting side project.)

I was reminded by a couple of recent current customer studies of some of the reasons Batch and CICS often don’t run alongside each other. Sometimes it’s logical "end of day" quiesce points (or "Positions") and sometimes it’s to ensure CICS and Batch don’t compete for resources. (More than half the customers I know have a higher Batch Window CPU Utilisation (if you squint at it) than that of their online day.) The "end of day" reason often shows up as CICS closing data sets and batch jobs opening them (and conversely at the other end of the night). I saw this in the study I’m currently finishing off. As many of you know I use the Life Of A Data Set (LOADS) technique – and this time I saw CICS regions as well as batch jobs. I think it would be useful to see how long after CICS comes down the first batch job processes any of the region’s data. And the same at the other end.

It’s much more difficult to see the DB2 objects – table spaces and index spaces – accessed by CICS regions and batch jobs. But then, as I said, overlap is more common here. It raises the interesting question of how an installation knows it’s safe to run CICS and the related batch concurrently. Maybe you can share your experiences. I think it has something to do with Applications people (shudder) designing things. πŸ™‚

The same applies to MQ queues, of course. And IMS is a whole other game.

So we can use SMF 30 to document uptime for CICS. I think it would be useful also to form a view of what the transaction profile looks like while the regions are up. This would probably be driven by CICS SMF 110 records, possibly pulling in correlated DB2 SMF 101 and MQ SMF 116 records. There are two reasons to do this:

  • Effective outages – where the region is up but work still can’t get done could be documented. (A healthy transaction rate suggests a healthy region.) Actually a spate of "unhappy ending" transactions might mean something – such as the loss of the database manager or some partner region.
  • It would be interesting to see how transactions peter out towards the end of the day (if they do) and perhaps "peter in" πŸ™‚ at the beginning. You might use this to tell you CICS could afford to be up for less time (to make room for the growing Batch). I’d prefer to think of it conversely: Justifying the need to keep regions up as late as you do and starting them as early as you do. Taken to its logical conclusion it might justify a project to make the CICS and Batch run concurrently.

By the way everything I’ve said above (other than the specific program name) applies to most of the other online application styles. I’m just reminded of CICS, as I say, by a couple of current customer engagements. Undoubtedly the next one will remind me of something else. πŸ™‚

DB2 Package-Level Statistics and Batch Tuning

(Originally posted 2011-12-05.)

I don’t know how many years it’s been since DB2 Version 8 was shipped but I’ve FINALLY added support for some really useful statistics that became available with that release.

As so often happens I was caused to open up my code because of some customer data that exposed a problem in it: The customer sent DB2 Version 8 SMF 101 Accounting Trace data that contained Unicode. In particular DB2 Package names were showing up as apparent garbage. Hexdumping some records showed this field to be Unicode.

The first step was to tolerate Unicode. In my case I translate it in the REXX I do my actual reporting in. (I could’ve done it in the assembler and maybe one day I will – but it makes the job more complex.) There is a field in the Product Section (QWHSFLAG) that has the value x’80’ if the record contains Unicode (and 0 if it doesn’t).

But this post isn’t really about Unicode. It isn’t about the longer names that are supported in Version 8, either. It’s about some nice "new" statistics you also get at the Package level. (And nothing significant has happened to the 101 record since Version 8.) As I had the code open I took the opportunity to exploit these new numbers. I’ve not written about themΒ  before so now is a good time to extol their virtues – despite the arrival of Versions 9Β  and 10.

So this post is about DB2 performance at the package or program level. "Program" would be the application code or not-specifically-DB2 term: An application program calling DB2 generally uses a package with the same name. I’ll use package in the rest of this post because it’s the DB2 term.

Buffer Pool Statistics

The Accounting Trace record has very nice buffer pool statistics at the Plan / Buffer Pool level. But the real problem for a batch job is "which program / DB2 package is driving the traffic?" We’ve always had the ability to say which packages the time was spent in and which components of response time for those packages are dominant. And indeed the major packages might be spending lots of time waiting for Synchronous Buffer Pool I/O or Read (or Write) Asynchronous I/O. (I see that quite often.)

What we didn’t know until Version 8 is how the buffer pools are performing for those top packages. So these statistics are really handy.

Note: There’s only one set of buffer pool statistics for each package. That is, you can’t tell which buffer pools are accessed by which package.

SQL Statistics

At the plan level we see the number of, for example, singleton selects, cursor opens, fetches under cursor, updates, inserts and deletes. So we might, for instance, gain some insight into why a batch job step is seeing a large amount of Synchronous Database I/O Time: Perhaps it’s because of a plethora of singleton selects so Prefetch doesn’t really happen.

What we couldn’t do, prior to Version 8, is see this at the package level. Now we can we’re able to find the package / program that’s behaving this way. So we stand a better chance of fixing it.

As an experiment I summed up counts of the different SQL statement at the package level and compared it to the field QPACSQLC (which was there at the package level long before Version 8). This field is the number of SQL statements. Usually they’re the same but in a significant proportion of cases the sum is less than QPACSQLC. One valid explanation is that the difference includes commits and aborts (which there aren’t statistics for at the package level). I bolded "includes" because this isn’t the whole explanation. If you take out the plan-level commit and abort counts (fields QWACCOMM and QWACABRT) you sometimes still have a discrepancy. I’ll have to research why this might be.

But, as I say, the reason for this post about old (but not obsolete) statistics is that I think y’all will find them really handy. And especially for batch steps and indeed any DB2 application where the transaction comprises multiple packages / programs.Β