Batch Capacity Planning, Part 1 – CPU

(Originally posted 2011-09-19.)

It’s been a week since the following was posted in IBM-MAIN: Batch Capacity Planning – BWATOOL? So far there’s been no reply. Though a little disappointed, I’m not surprised. "Disappointed" as I was looking for a good debate (even though it wasn’t me who asked the question). "Not surprised" as I think the subject of Batch Capacity Planning is a tough one. The original post prompted me to think about posting on the subject. I think I’ll do it in two parts:

  1. CPU
  2. Memory

An obvious place to start is by comparing and contrasting Batch with Online, from the Capacity Planning point of view. This, of course, builds on the Performance / Window perspective:

  • Online comprises discrete (maybe discreet) 🙂 pieces of work, seemingly unconnected. Batch work – at least the stuff we tend to care about – is a network of inter-related pieces of work.
  • Online pieces of work (transactions) are brief relative to any summarisation interval you’re likely to use. Batch jobs, while many are very short, are often long compared to an interval. Although I’ve seen windows where the jobs are at most 15 minutes most Batch has key jobs in it that are much longer than this (the standard RMF) interval.
  • Online transactions (and their kin) are "scheduled" by being requested. Production Batch tends to be kicked off by a scheduler. (The "and their kin" parenthetic comment refers to the fact that many workloads are transaction-like, such as many styles of DDF requests.)

Those contrasts aren’t exhaustive but they are enough. We’ll use them to inform the rest of this post.

But there is a similarity that’s worth articulating: For both Batch and Online "enough CPU" refers to "what gets the job done": If Online work fails to meet Service Level Agreements / Expectations / Pious Hopes 🙂 or whatever you conclude something has to be done. Similarly, if important Batch fails to meet its business goals there’s pressure to do something. (This post isn’t going to go into the business drivers or the shape of SLAs.)

When I look at the CPU Utilisation for a Batch Window I typically see a huge amount of variability, both within the night and from night to night. This, I surmise, is caused by the "big lumps" characteristic above. And if you try to figure out which job caused a spike it’s hard to do automatically – because of the "interval straddling" characteristic above. But usually it’s fairly obvious – if the number of jobs running is not too large – which job is likely to have caused a spike.

 It also helps if you have a decent WLM Batch service class (and, hopefully, reporting class) scheme: You can identify which service class caused the spike. And thence the list of candidate jobs could be shorter.

WLM setup helps in another way: Assuming you have a sensible hierarchy of Batch service classes you can establish whether the supposedly-more-important jobs’ service classes are experiencing delays for CPU (from the WLM delay samples in RMF). With this view you can take a "squint at the picture" * approach – smoothing out the spikes to some degree. In fact summarising over a longer interval than you might normally could be useful.

I think you have to accept that some degree of delay is inevitable at times with spiky work like Batch. Even for the "top dog" Batch service classes. The question is "how much?" If you calculate WLM velocity for these service classes over a long enough interval and the work of the window is just completed, maybe that’s a useful metric and threshold. When the velocity drops below a certain level the window’s work might just fail to get done in time.

I appreciate the previous paragraph is a little vague: It’s trying to impart an approach to learning how your Batch works – from the CPU perspective. The "inevitable at times" phrase might be a little controversial: Certainly if your Online day drives the CPU capacity requirement you stand less of a chance of seeing CPU delays of any note in the Batch. But for many installations that’s not true: The Batch drives the CPU requirement (or, in some cases, drives the Rolling 4-Hour Average and hence the software bill).

I haven’t used the "job network" characteristic in this post yet. So here are a couple of areas I think it plays in:

  • It stands to move the spikes around: When a job gets elongated or delayed things move in the window. That could be the job itself being extended or downstream jobs being delayed.
  • Growth is as likely to express itself in  delays to jobs’ starting and completion as in increasing CPU requirements. Growth is mainly about "more data", though Development might add function so each input datum "gets more attention". 🙂

These two are related, I think. And they’re both about topology.

A couple of other things:

  • Given the complexity of Batch it’s very difficult to predict what the effect of tuning is – on either run times or capacity requirements.
  • It’s also very difficult to figure out how the "more engines or faster?" debate plays. "More and faster" is a clear winner (or at least non-loser). Some parts of the night will favour more. Some others will favour faster. Again the "squint at it" * technique is reasonable: If generally the answer is more, though sometimes it’s faster, probably more wins out. But apply common sense: Optimisation in one place causing de-optimisation in another requires care.

 I’ll admit this whole area is a tough one. I’d be interested in what customers do for Batch Capacity Planning – or indeed whether it drives their overall plan.

And soon I’ll write about Memory from the Batch Capacity Planning perspective.


* When I say "squint at it" I mean "use a technique that takes the spiky detail out, leaving an overall (if blurred) picture. I’ve used the term for many years. People don’t look at me oddly when I say it so I assume they know what I mean. 🙂  

zAAP and zIIP Revisited (At Last)

(Originally posted 2011-09-14.)

It’s been an awfully long time since I wrote When Good Work Doesn’t Go To zIIP / zAAP Heaven. And too long since I posted anything at all. 😦 In fact I had a bunch of posts in my brain until this morning when a customer asked me a question which turned into this blog post. (Those posts are still in my brain and will probably see the light of day eventually, including one based on a customer question.)

They wanted to know whether there was an In-And-Ready count for zIIP or zAAP. I don’t see such a thing. But what I do see is, in my opinion, much better. I’ve also presented a chart in my "Much Ado About CPU" presentation on the subject for some time now. I’m surprised I haven’t blogged about it already. So here it is, while it’s still useful: 

Take a look at this graph:

It’s for a single WLM Service Class period across a 24 hour period. On the vertical axis we have a couple of mixed types:

  • A stack of the CPU components for the service class period, as a percent of an engine. These are the bars.
  • A yellow line representing the percentage of WLM samples that represent zIIP delays.

In this example it’s a relatively benign case. Here’s how I read it:

  • The red bar values are "GCP on GCP" – the work that was never eligible for zIIP. It’s normal for it to run this way.
  • The blue bar values are "zIIP on zIIP" – the work that was eligible for zIIP that actually ran on zIIPs. This is goodness.
  • The green bar values are "zIIP on GCP" – the work that while eligible for zIIP ran on GCPs. This is what you’d like to minimise.
  • Because almost half ran on GCP I conclude this is DDF work. (In fact it running on a zIIP before zAAP-on-zIIP and the name of the service class confirm this. The regularity of the pattern and split also corroborate it.)
  • At times when the delay samples tick up so does the "zIIP on GCP". This is the key correlation.
  • In fact the "zIIP on GCP" portion isn’t all that bad. I’ve seen worse and I’ve seen a little better.

This is a standard part of my reporting. Hence I’m in a position to say "I’ve seen worse and I’ve seen a little better." I have experience these days. 🙂

 

Some observations:

  • It’s probably fair to use normalised CPU – particularly as there are many machines where the specialty engines run at full speed and the general purpose ones don’t.
  • It’s probably a good idea to add in the "Delay for GCP" sample percentage. I was sensitised to this by a customer both of whose pools were – for the service class period in question – showing serious delay samples.
  • In general this chart has both zIIP and zAAP on it for the same service class period. The technique works for zAAPs, zIIPs and zIIP on zAAP just the same.
  • Talking of zAAP on zIIP: I expect the numbers to all look like zIIP: There are no specific zAAP on zIIP metrics.
  • This technique allows you to evaluate things like IFAHONORPRIORITY and IFACROSSOVER as they work out at the service class period level.
  • The PROJECTCPU mechanism only works for workloads that are already running. For example, turning on IPSEC is a fresh workload: It won’t show up until you run it (whether on a GCP or a specialty engine).
  • If an exploiter changes how it behaves (for example DB2 DDF with PM12256) you’ll see some clues in this chart. I say "some" because in that particular APAR the variability in outcome at the thread instance level is not going to show up here. It might show up in DB2 Accounting Trace (SMF 101) if you go looking for variability at the individual SMF record level.

I think the graph works well (even if the colours etc don’t)  and I think it’s a chart you can replicate and build on (including the customer who asked the original question).

One final (meta) point:  I hope that if you ask me a question that’s of wider interest you won’t mind me posting the answer (perhaps extended) as a blog post like this. Of course I’ll shoot you the link. And, as you’ve seen in the past, of course I’ll avoid posting your data – unless you OK me doing so.

Setting Arbitrary Variables In A REXX Procedure That Persist After It Completes

(Originally posted 2011-08-23.)

Over the weekend I decided to try my hand at writing a JSON parser in REXX. I developed the technique outlined in this post to enable me to do that. I think it’s new. Certainly I’ve not seen it before.

(I don’t want at this point to get into a discussion about why JSON parsing might be a useful thing to do. Suffice it to say it is possibly going to be second only to XML as a common data format in the future.)

Consider the following code:

/* rexx */ 
stem="mystem" 
call setvars stem 
say "MYSTEM.A="mystem.a 
say "MYSTEM.B="mystem.b 
exit 0 
                             
setvars: 
parse arg stem 
interpret setvars2(stem) 
return 
                             
setvars2: procedure 
parse arg stem 
return stem".A=1;"stem".B=2;"

Here’s the problem it’s illustrating a solution to: I want to call a routine that sets arbitrary REXX variables but otherwise doesn’t create any side effects.

You’ll notice the “procedure” in the setvars2 routine that avoids the side effects. That’s pretty standard, actually. But it doesn’t allow any variables to be set. And I need them set.

So I return a string from setvars2 to setvars containing assignment statements. In this case the string is “MYSTEM.A=1;MYSTEM.B=2;” and it’s immediately executed – using the interpret statement.

The setvars routine (as you’ve probably guessed) is a wrapper. While it doesn’t use procedure – as I want the variables it sets to persist – it’s pretty clean: To the caller it appears clean, anyhow – as they oughtn’t to notice setvars2 that’s doing all the work.

There’s one fly in the ointment: The variable called “stem” gets set. I consider that a minor untidiness. If you, dear reader, know of a way of removing it I’d be interested in hearing it.

So it is possible to set variables in a routine that’s otherwise protected by procedure. You’ll notice that the variables that can be set are bounded by the stem passed on.

Seasoned REXX programmers are likely to think that expose solves the problem. Actually it doesn’t 😦 : I simply don’t know what the variable names will be and expose won’t take a string I could build. I thought about that and tested it.

You’re probably still wondering about my motivation: In the actual JSON parsing code you have to handle nested structures. So my equivalent of setvars2 calls itself recursively, passing back up a string containing assignment statements, separated by semicolons. These strings are appended to each other and the interpret in the equivalent of setvars indeed executes them.

We’re talking about resulting strings like:

"s1.a=1;s1.b.c='Hello';s1.b.centre.x=-5;s1.b.centre.y=22;"

which would come from a JSON string:

{
  "a":1,
  "b":
  {
    "c":"Hello",
    "centre":
    {
      "x":-5,
      "y":22
    }
  }
}

One of the issues here is knowing which variables got set – as we don’t know the JSON in advance (for certain). The way I can see that being handled is to set a “dictionary” variable containing the names we gathered up along the way. And then we get into the question of JSON Schemae – which I don’t intend to discuss at all here.

And the reason for mentioning all this is that I can see the technique has a wide range of applications. For me it certainly gets around a thorny problem involving recursion.

And no, it wasn’t the highlight of my weekend: Probably that was Walkway Over The Hudson which I’d wanted to do since long before it opened 2 years ago.

Another Graph I’m NOT Going To Share With You – Batch Window Reduction Expectations

(Originally posted 2011-08-12.)

In WLM Velocity – "Rhetorical Devices Are Us" I mentioned a graph I wasn’t going to publish – essentially to protect a customer. In this post I’m again going to describe a graph I have (at least in my head) without publishing it. (And for essentially the same reasons.) I hope you find it useful, however:
 
I’ve been acquainted with a lot of customer Batch Window Reduction projects, especially in the past two years. When I hear of a new one I "plot" the situation on the "graph in my head". So, if I tell you that what you’re trying to do is likely to succeed (or conversely that it isn’t) that’s (mostly) where it’s coming from.
 
So, about that graph:
 
  • On the x axis I plot the degree of scale up you’re trying to achieve. So, I might talk about "2X" or "1.3X".
  • On the y axis I plot how much time you have to achieve that in. Which might be as little as "by next month end" (no really!) or "within the next 2 years".
 
You can probably guess what’s coming next:
 
If you are trying to achieve a big scale up (say 5X, to quote a real recent case) in not much time (say 3 months, thankfully not the same case but nonetheless a real value) then I’d classify that as very high risk. (That’s the bottom right hand corner of the graph.)
 
On the other hand 1.3X in 2 years is very low risk and is in the top left hand corner of the graph. (I’ve never seen anyone that lucky.)
 
So that’s the graph – just another rhetorical device. But there are some wrinkles. I’ll list a few – and you can probably think of a few more:
 
  • What if we have "way points"? Say 1.5X in 6 months and 2X in a year?
  • By which metric is it, say, 2X?
  • This says nothing about cost. Zero cost probably means zero speed up. Infinite cost might make the batch go a lot faster.
  • This doesn’t speak to efforts already undertaken. Obviously a thoroughly-well-tuned application is unlikely to speed up as much as an untuned one.
For these reasons I wouldn’t take the graph too seriously and I don’t entirely rely on it. It’s also a part of why I’m leery of committing to precise speed-up numbers.
 
Being a Performance person I hope any fuzz in my language reflects the fuzz in the situation, not a general unwillingness to commit. If you think I’m doing the latter be sure to tell me so.
 
In the meantime I hope you find this rhetorical device useful.

Back From Vacation And Raring To Go – To Poughkeepsie

(Originally posted 2011-08-11.)

Usually when I go away on holiday I bring something back with me. Often in the form of fresh ideas. This year it’s been such a hectic one that all I did was to flake out. So no new ideas this time. Perhaps that’s a good thing, perhaps not. 🙂

But I think I did achieve something: Mental decluttering. So I can, for example, look at stuff I was working on with a fresh take.

And the timing is actually pretty good for that: On Sunday I fly to Poughkeepsie to begin a four-week residency. We’ll be writing a Redbook on "Batch Modernization", following on from Batch Modernization on z/OS, which people seem to have rather liked. 🙂

I actually don’t know who will be on the team: I expect a mixture of the previous team (any of whom I’d be glad to be working with again) and new people (pleased to meet you). 🙂

I also don’t know what we’re going to write about. So I really do start with a "clean mind". 🙂 Or at least, I hope, an open one.

I think some of what I’ve talked about in recent posts could be useful – if not immediately reused – in the Redbook. I’ve also a few bits and pieces in my mind. But it’s a team effort to define shape and content. (And one of my ideas is that we look at each other’s stuff more this time around – to provide a different perspective.)

Now, I think there are two things that concern you, dear reader :-) :

  • A question: Are there topics in the Batch Modernisation realm you’d particularly like to see covered? No promises, but I am interested…
  • A hope: I’d like to think I could take some extracts from what we’re writing and post them here. Again, no promises as the rest of the team might not be happy with these "teasers".

I should point out that it’s not MY residency: My friend Alex Louwe-Kooijmans is running it. And, further, it’s a team effort. But I’m raring to go, recharged from holiday, and hoping to share what I can with you.

(And actually I did experiment with one thing while away: Programming Mac OS X – both Applescript and with Objective C. It’s the first time I’ve really had the chance to get to know my Macbook Pro.)

Another Neat Piece Of Algebra – Series Summation

(Originally posted 2011-07-19.)

Here’s another neat piece of algebra: A technique for summing series.

You know what b – a + c – b + d – c is. Right?

Suppose I were to write the same sum as:

(b – a) +

(c – b) +

(d – c)

The answer is still d – a. Right?

Now, suppose I re-label with s0 = a, s1 = b, s2 = c and s3 = d. We end up with:

(s1 – s0) +

(s2 – s1) +

(s3 – s2)

This is actually pretty scalable terminology as you can write sr for any arbitrary value of r. And that’s one of the strengths of algebra: generalisation.

So let’s do that up to n:

(s1 – s0) +

(s2 – s1) +

(sn – sn-1)

which is, of course, sn – s0.

But what has that got to do with summing series?

If we can replace each (sr – sr-1) by a single term ur you may see the relevance…

u1 + u2 + … + un = sn – s0.

The series summation boils down to a "simple" subtraction. The trick is to find these s terms, given the u terms. Let’s try it with an example.

 



Summing The Integers

 

This is the series 1, 2, 3, … , n.

The r’th u term is just r. ur = r. So we now have to find the sr term. Remember sr – sr-1 has to equal r.

Try sr = r (r + 1).

Then sr-1 = (r – 1) r   or   r (r – 1).

So sr – sr-1 = [(r + 1) – (r – 1)] r   or   2r. Not quite what we wanted. But we know – dividing by the factor of 2 – we should’ve guessed sr = ½ r (r + 1).

So sn – s0 = ½ n (n + 1) – ½ 0 (0 + 1) = ½ n (n + 1) – 0 = ½ n (n + 1).

The sum of the first n integers being ½ n (n + 1) is a well-known result. Admittedly it could’ve been done another way. But it’s simple enough to show the method.

 


Another Example – Summing The Squares Of The Integers

 

This is the series 1, 4, 9, … , n² .

In this case we need to do something that will appear slightly perverse:

Rewrite r² as r ( r + 1) – r.

If you can split each of the terms in a series into two terms you can sum these sub terms. I just did the split. We already know how to sum the "r" portion. It’s ½ n (n + 1). So we need to sum the r (r+1) portion and subtract ½ n (n+1) from the result.

Try sr = r (r + 1) (r + 2).

Again we need to find sr-1.

It’s:

(r – 1) r (r + 1)

or, rearranging,

r (r+1) (r-1)

So

sr – sr-1 = [(r + 2) – (r – 1)] r (r + 1) or 3r (r + 1).

This is 3 times what we want so we should’ve guessed sr = 1/3 r (r + 1) (r + 2).

So this portion of the sum is 1/3 n (n + 1) (n + 2) – 1/3 0 (0 + 1)(0 + 2) or 1/3 n (n + 1) (n + 2).

But we need to subtract ½ n (n + 1) from this:

1/3 n (n + 1) (n +2) – ½ n (n + 1) = 2/6 n (n + 1) (n +2) – 3/6 n (n + 1)

or

1/6 n (n + 1) [ 2 (n + 2) – 3] = 1/6 n (n + 1)( 2 n + 1) .

If you try it for a few values you’ll see it’s right. This isn’t such a well known result as for the sum of the integers.

 


 

I’m conscious there’s been some fiddliness here – which is where I normally fall down. 😦

But I think the "sum a series by converting it to a single subtraction" trick is a neat one – which is why I share it with you.

An Experiment With Job Naming Conventions

(Originally posted 2011-07-19.)

It may surprise you to know I hate asking questions to which I already know the answers. :-) And I hate even more "leaving understanding on the table". Let me put it more positively: I love it when I can glean new insights into existing data. This post is about precisely that: An experiment in gleaning extra understanding…
 
In Batch Architecture, Part Zero and follow-on posts I talked about gleaning how an installation’s batch applications fit together. I’ll admit that part of it was a little sketchy and I’ve had the opportunity since then to look at a number of customer batch environments. I really don’t much like the part where I ask the customer "what’s your batch naming convention"? So I wrote some experimental code and tested it with one of these recent sets of data…
 
My raw data in this is SMF 30 Job-End records, processed into a database in my usual way. (And you, too, could do the same – and everything else that’s in this post.)

Remember I’m looking for patterns in 8-character tokens, and about 100,000 of them. The latter may be an under- or an over-estimate for you. The former is fixed. (And this technique might work with other bounded-size tokens such as DB2 Accounting Trace Correlation IDs or CICS region names.)

Here’s the process my code follows:

  1. Discern some masks from a pass over the data. (More about this towards the end of the post – but it is the first step.)
  2. Apply these masks to all the jobs and see which masks fit. (I’ll tackle this first as it explains why we need to do Step 1.)
 

Do These Jobs Match This Mask?

In this post a mask is a string of characters (for example "AAA999AA") against which each job name is tested. The "A" denotes "any alphabetic character in this position" and the "9" denotes "any numeric character in this position". So, in this example, a match would be a job name with the first three characters alphabetic, the next three numeric and the final two alphabetic.

(It’s perfectly reasonable to complicate things by allowing more than just "A" and "9". Perhaps "$" for non-alphanumeric and "*" or "?" as wildcards. I really don’t think that level of sophistication is necessary for this prototype – and Regular Expressions are probably overkill*.)
 
Because I knew the test data I used the espoused naming convention for the customer: "AAA999AA" is indeed the mask for this. My code shows that 86% of all batch jobs match this naming convention. So what about the other 14%? 🙂 Maybe that’s a metric: percent_jobs_matching_espoused_naming_convention. :-)
 
I could’ve stopped there but I thought it useful to analyse the three-character "AAA" piece of the mask: There were 35 different values. Sorting these by occurrence descending I see 11 with over 100 occurrences (the top one having 732). These could be suites (or applications, if you prefer). This I’d be happy to share with a customer. It would enable the conversation to start somewhere more useful than "what is your naming convention?"
 
But, you’ll note, that’s one mask ("AAA999AA") that was already handed to me. Nice but not enough. I still think this "leaves understanding on the table".
 


 

How Do I Generate The Masks?

As I said, I think I can teach my code to do better than that. In fact I think I did…
 
With 8-character masks where each mask position can be in one of two states ("A" or "9") there are 256 potential masks (and that’s probably only 128 as I think the first position will have to be "A" – not that I’ve coded with that assumption). The point is there isn’t much potential for an explosion.
 
I glean the masks the following way. I run through all the job names, one character at a time:

  • If the character present in, say, more than 90% of the job names is a letter I add "A" to any (partial) masks already generated.
  • If the character is more than 90% of the time a number I add "9" to any partial masks.
  • If not I create two sets of masks – one with an "A" on the end and one with the "9" on the end.

In this test I generated four masks: "AAA999AA", "AAAA99AA", "AAA9A9AA" and "AAAAA9AA". All the masks start with "AAA" and end with "9AA". The doubt is in the middle where "99", "A9", "9A" and "AA" got generated.

If I drop the threshold from 90% to 80% I only get "AAA999AA" so maybe that is a good naming convention after all. (In fact the middle characters are 87% and 88% numeric, respectively. And the sixth character is numeric 91% of the time – so it scraped through.)
 
As I said, my initial testing of the mask-matching used "AAA999AA" because the customer had indicated that was their convention. So my code allows you to specify masks and then adds the automatically-generated ones to it.
 


Conclusion

I think the experiment worked well. I can see cases where the code needs enhancing. I can see cases where it mightn’t be perfect. But I do think this code worth running (and tweaking) at the beginning of every relevant engagement.
 


* I’m doing my programming in REXX – which doesn’t even have regular expressions. It might be nice to write a function package that did it. A challenge for someone? Anyone? 🙂

Multiline Message Sifting With DFSORT

(Originally posted 2011-07-17.)

Frank Yaeger of DFSORT Development suggested I pass this tip along to y’all. It’s his solution to a problem set by Brian Peterson of UnitedHealth Group…

In z/OS Release 12 two new messages were introduced: IEF032I and IEF033I replace IEF374I and IEF376I. The older messages were single-line step- and job-end messages. The new ones are their multiple-line analogues: IEF032I is 3 lines and IEF033I is 2 lines.

The problem is how to sift these messages out in a program.

Here’s Frank’s solution:

//S1 EXEC PGM=SORT
//SYSOUT DD SYSOUT=*
//SORTIN DD DSN=...  input file (FBA/133)
//SORTOUT DD DSN=...  output file (FBA/133)
//SYSIN DD *
  OPTION COPY                                                 
  INREC IFTHEN=(WHEN=GROUP,BEGIN=(2,7,CH,EQ,C'IEF032I'), <1>      
    RECORDS=3,PUSH=(134:ID=1)),                               
   IFTHEN=(WHEN=GROUP,BEGIN=(2,7,CH,EQ,C'IEF033I'), <2>           
    RECORDS=2,PUSH=(134:ID=1))                                
  OUTFIL INCLUDE=(134,1,CH,NE,C' '),BUILD=(1,133)             
/

It uses DFSORT’s IFTHEN WHEN=GROUP to form groups. Overall the trick is to assign a group number to those records that are part of the message (placing it in position 134) and leave the records which aren’t with a blank in position 134.

The OUTFIL statement keeps only those records without a blank in position 134, truncating the output to 133 bytes.

So, how does it distinguish between those records that are part of an IEF032I or IEF033I message? The INREC statement uses a pair of IFTHEN clauses in a short "pipeline":

  • IFTHEN clause <1> begins a group when it encounters "IEF032I" in position 2 (after the ASA control character). For three records beginning with this one a group number is assigned (in position 134). It’s 1 byte long – specified with "PUSH=(134:ID=1)". It doesn’t matter what the group number is, so long as it’s there for the OUTFIL statement.
  • IFTHEN clause <2> does the same for the "IEF033I" message. This time it’s two records beginning with the "IEF033I" line.

Looking back I see I used the term "pipeline" for IFTHEN in Unknown Unknowns in 2009. This example incorporates a two-stage pipeline: Any records not satisfying the BEGIN= clause in <1> (the first stage) are passed to the second stage (marked <2>). But no records that satisfy the clause are passed on. (If we wanted them to we could code "HIT=NEXT" – but that wouldn’t be useful here.)


The code I’ve shown above is just what Brian needed – and I haven’t altered it from what Frank sent me.

I think there are three fairly obvious twists to this:

  • If you wanted to you could decode the multi-line message – perhaps with PARSE.
  • If you wanted to support the old messages (IEF374I and IEF376I) with more IFTHEN clauses. You might do this if had multiple LPARs – not all at or above z/OS R.12 – to support. (Or maybe you’re a software vendor.)
  • You could support multiple messages – not just the two versions mentioned here – and you might use OUTFIL to route them to multiple destinations. (I think you’d need another flag for that.)

But Frank’s example is a very good one and I’m glad he sent it to me. This is a sort of "guest post" though I did all the writing. I wonder if there are topics you would like Frank to actually "guest post" on. I can put them to him.

Hello, I’m Martin And I’m An Algebraic :-)

(Originally posted 2011-07-09.)

If you’re sat next to me on a plane you’ll probably notice at take off and landing I do algebra puzzles. You may not have heard of the term "algebra puzzles" before and perhaps think the juxtaposition of the two words to be odd, but I think it apt…

(You may also think this whole post to be showing off, but that’s a risk I take in sharing a passion I have.) 🙂

A classic problem with take offs and landings is what to do given you’re not allowed to use electronic equipment. I’ll readily agree that staring out the window is a good one – which is why I prefer a window seat. I love staring out the window. I love maps – and to me looking out of an airplane window brings maps to life. And figuring out what I’m seeing is another great puzzle. But sometimes there’s nothing to see. So what do you do?

I started by taking puzzle books with me. I’ve done Sudoku (but not recently), Kakuro, Futoshiki, Hashi, Kenken and any number of others. I enjoy them but each one lacks variety. (And I’m disappointed that by far and away the most common puzzle books are Sudoku.)

But I find the best puzzles of all are algebra problems. I still have a copy of my "high school" Further Mathematics textbook. I don’t know why, I just do. 🙂

I actually think it’s the elegance of expression and the neatness of the right shortcut that appeal to me. As I’ve said many times I’m a sucker for ingenuity. Below is an example of a neat shortcut that I’d like to share with you. I hope you’ll see what I mean.

One of the nice things about mathematics in general is that you’re perpetually "standing on the shoulders of giants". Some of them well known (Newton, Leibnitz, Euclid, Gauss, etc) but many are anonymous. In the example below I’ve no idea who thought of the shortcut first. (I’m just pleased I understand it and can see its applicability.)


A Simple Example Of Elegance

Problem: Solve (x – 3)² – (x + 2)² = 0

It looks like a difficult puzzle to solve. Of course if it were I wouldn’t be offering it here. 🙂 You could multiply everything out and gather terms but that’s horrid. Thankfully, there is a more elegant way:

Observe x² – y² = (x – y) (x + y) . (Check it if you don’t believe me!)

If you substitute a for (x – 3) and b for (x + 2) you get:

a² – b² which, of course, can be rewritten as (a – b) (a + b) .

I think you’ll agree working out what (a – b) and (a + b) are is easy:

a – b = (x – 3) – (x + 2) or -5

a + b = (x – 3) + (x + 2) or 2x – 1

Multiply them together and you get:

-5 × (2x – 1) = 5 – 10x

So (x – 3)² – (x + 2)² = 5 – 10x which = 0, as the original problem stated.

If 5 – 10x = 0 then 10x = 5 and so x = ½.


 

See, that wasn’t so hard, was it? I think people think mathematics is hard. I don’t think algebra is hard. I do thing topology is hard – because of the abstractness of the concepts. I do think proving things is hard – because of the need to not miss any loose ends and to know whether you’ve actually proved anything. But algebra is, to me, pure puzzle solving. And elegance is important: In the above example I could quite easily have made a mistake if I’d not known the trick. With the trick I’m much less likely to.

Now someone will probably come along and point out a few things about the example, including a further trick. If they do I’ll be delighted. This "old dog" loves learning new tricks. 🙂 And if I am sat next to you on the plane at least I won’t be muttering to myself as I manipulate those symbols. 🙂

10,000 Hours Doing WHAT?

(Originally posted 2011-07-07.)

It’s a popular suggestion that what separates the truly exceptional person from the rest of us is 10,000 hours of "practice". In book form I’ve seen it twice – in Matthew Syed’s "Bounce" and in Malcolm Gladwell’s "Outliers". Actually, to be fair, Matthew acknowledges his original source so that’s actually only one distinct source. (Life Lesson aside: Trace ideas and "facts" back to see if they came from one place or are truly corroborated.)

The suggestion is that there’s no such thing as innate talent and that all that matters is practice – 10,000 hours of it. And this is often repeated now in public folklore.

I find this assertion problematic for a number of reasons – though I’m prepared to admit I could be wrong.

For a start I find it very hard to believe we all respond the same way to our experiences, that we have the same physiology and brain "wiring".

Secondly – and this is where the title of the post comes in – 10,000 hours of what? Actually the suggestion is that it is useful practice. The issue for me is that the usefulness of the practice is highly variable – depending on who we are, how motivated we are and how tired we are (and maybe many other variables besides).

Thirdly, there is what I call the categorisation problem: Take the example of someone who is a generalist in, say, Information Technology. They could easily gain 10,000 hours of experience, perhaps good experience, spread across their whole domain. Does this make them an expert? If they spent the whole 10,000 hours in a narrower area what kind of an expert does that make them? And if they spent two chunks of 5,000 hours in two areas are they not an expert in either? What if those two areas were abutting? What if they weren’t?

That previous paragraph had a lot of questions in it. Some are easy to answer, some less so. Maybe – and here I hope is the relevance of this whole "10,000 hours" idea – these sorts of questions allow one to "evaluate" one’s career. I put "evaluate" in quotes because I don’t mean this as a scorecard: There are very many valid paths through life. But, for example, discovering your 10,000 hours have been spent scattered across a wide range of topics might tell you you’re a generalist. But then you probably knew that. 🙂

A more difficult case is what to do when you discover you’ve spent 10,000 hours in the same area. With a low boredom threshold like mine there’s a premium on convincing yourself there’s been some diversity. 🙂 But has there really? Or the converse: 10,000 hours concentrated in one area, but is it really one area?

In my case – and this post isn’t really about me – I’ve convinced myself 🙂 of three things: That there’s plenty of variety, that the technology keeps evolving at a dizzying pace, and that my role has in any case morphed over time. Actually, I think a lot of us feel that way.  

But does all this change keep resetting the counter to zero hours? I’d maintain it didn’t: The way I learn and (I think) the way I incorporate situations into my experience base is accretive (adding on around the edges of what I already know).  So I don’t really think the counter ever resets: We just start new counters occasionally. But I wouldn’t count some of the new technology I dabble in as "start at zero" – despite how much it sounds like babbling. 🙂

I enjoy being in the babble phase. But, like most people, I worry about the quality of the work I produce in that phase. I particularly enjoy when I spot the experience beginning to build. That "10,000 hour"idea just might be motivational. And now for a gratuitous Metallica lyric. 🙂

"Trust I seek and I find in you
Every day for us something new
Open mind for a different view
and nothing else matters"

In the spirit of "10,000 hours" don’t you think that’s a great verse? Anyhow, keep logging those hours. Who’s to say "it’s all been a waste of time"? 🙂

| Minor edit 27 October to fix "whose" that should’ve been "who’s". Dunno how that crept in there. 😦