XML, XSLT and DFSORT, Part Three – Multiple XML Input Files

(Originally posted 2011-05-22.)

While I was putting together the original three posts in this series a number of thoughts struck me, amongst which two really cried out for further investigation:

  1. I don't know how your XML data arrives on z/OS but quite a lot of scenarios don't have the data all as one document (file).
  2. XSLT looks complex – particularly if recursion does your head in. πŸ™‚

Thought 2 I'll deal with in a different post. This post relates to thought 1.

 

In the example I've given there are three item elements, representing three transactions. I'm not sure that's entirely realistic:

Certainly there will be many times (for example configuration files) where everything is in the one file. But consider the following scenario:

XML documents arrive in a directory, each representing a single transaction, or maybe a batch of them. In the rest of this post it doesn't matter whether there's more than one transaction in a file, only that there will be multiple files overall. In the previous three posts I've talked about processing a single file with XSLT and passing the results to DFSORT (in a manner the latter can work with). The technique outlined won't (unaltered) process more than one file at a time.

I'd like one DFSORT processing run to handle multiple input XML files. Perhaps you run a batch job every hour to process all the transactions that arrived as XML files, or perhaps daily. The rest of this post shows you one way of doing this. It's a relatively small change to the XSLT stylesheet.

Here are the three transactions, as if they arrived in separate files:

Transaction File 1
<?xml version="1.0"?>
<mydoc>
  <greeting level="h1">
    Hello World!
  </greeting>
  <stuff>
    <item a="1">
      <row>One</row>
    </item>
  </stuff>
</mydoc>

Transaction File 2
<?xml version="1.0"?>
<mydoc>
  <greeting level="h1">
    Hello World!
  </greeting>
  <stuff>
    <item
      a="12">
      <row>Two</row>
    </item>
  </stuff>
</mydoc>


Transaction File 3
<?xml version="1.0"?>
<mydoc>
  <greeting level="h1">                                                        
    Hello World!
  </greeting>
  <stuff>
    <item a="903">
      <row>
      Three
      </row>
    </item>
  </stuff>
</mydoc>

XSLT can't directly process a list of files concatenated together. But you can do it if you can create another file. Here's an example:

Transaction Reference File
<?xml version="1.0"?>
<transactions>
  <transaction filename="txn0001.xml"/>                                        
  <transaction filename="txn0002.xml"/>
  <transaction filename="txn0003.xml"/>
</transactions>

If you can create such a file – perhaps by scanning an "incoming transaction file" directory – you can easily coax XSLT into processing the set of files. Here's a stylesheet that can do it:

XSLT Template Using The document() Function
<?xml version="1.0"?>
<xsl:stylesheet version="2.0"
  xmlns:xsl="http://www.w3.org/1999/XSL/Transform">

  <xsl:output method="xml" encoding="IBM-1047" indent="no"
    omit-xml-declaration="yes"/>

  <xsl:template match="/">
    <xsl:for-each select="transactions/transaction">                    1  
      <xsl:apply-templates select="document(@filename)/mydoc/stuff"/>   2      
    </xsl:for-each>
  </xsl:template>

  <xsl:template match="item">
    <xsl:value-of select="normalize-space(row)"/>
    <xsl:text>,</xsl:text>
    <xsl:value-of select="format-number(@a,'0000')"/>
  </xsl:template>
</xsl:stylesheet>

There are two things to note in this stylesheet:

  1. When you run the transaction reference file through XSLT with this stylesheet this line causes each transaction element to be visited.
  2. The filename attribute of each transaction element is used to pick up a transaction file (which might have one or might have more item elements in).

This "indirection through a transaction file" technique is very powerful.

In practice you might have an "inbound transaction XML" directory that you scan with a program that creates the transaction reference file, invokes (eg) Saxon and then invokes DFSORT, finally deleting all the successfully-processed transaction files. I say "(eg)" because nothing in this revised stylesheet requires XSLT 2.0 and so Saxon isn't the only choice.

I think the challenge in this is knowing when the transactions have been successfully processed and so the inbound files can be deleted. You'd have the same problem if – instead of creating a transaction reference file – you created one large XML file from inbound files. (In fact this is easier.)

Anyone feel like – in any z/OS-supported language – writing something to scan a directory for XML files and create a transaction reference file like the one above from their names?

Who Are You And What Have You Done With My Readers? :-)

(Originally posted 2011-05-20.)

I’ve done a little analysis of hits on recent blog posts. I wonder what you make of it:

Looking at this pie chart slices start at the top and go anticlockwise. Reading the legend is from latest to oldest, left to right, wrapping appropriately.

While the blog is called "Mainframe Performance Topics" I have lots of other interests. It’s interesting to see which posts have gained the most hits. It’s also fun to speculate on how these hits came about. I recognise hits don’t translate into reads and certainly don’t reflect what the readers thought of the post. Nonetheless:

  • The most popular topics are about Android and HTML5. I don’t think this is my usual readership. πŸ™‚ I think they flew in via web search. πŸ™‚
  • The multi-part posts (Batch Architecture, Vienna Conference, and XSLT / DFSORT) seem to do reasonably well. I’m guessing that people who read one part read the others.

I’m not particularly worried about numbers of hits, actually. I write because I think I have something to say. Also because it encourages me to learn from doing the research.

Although I can readily recreate this chart with additional data I’m not planning to do so.

But still, I think this chart is interesting. I hope you do, too.

XML, XSLT and DFSORT, Part Two – DFSORT

(Originally posted 2011-05-20.)

Following on from this post and this one, this post discusses the DFSORT piece.

The DFSORT code in this post parses the Comma-Separated Variable (CSV) file produced by XSLT processing. In this simple example it merely produces a flat file report, but the post has a few additional details you might find valuable.

First, here’s the SORTIN DD JCL statement. It’s not like a regular sequential file statement as it has to access the zFS file system we wrote the data to with Saxon:

JCL SORTIN DD Statement
//SORTIN    DD  PATHOPTS=(ORDONLY),RECFM=VB,LRECL=255,BLKSIZE=32760,  
//          PATH='/u/userzfs/myuserid/testXSL.txt',FILEDATA=TEXT 

Of particular note in this DD statement is the record format (VB), the logical record length (255) and the block size (32760). This is definitely VB data. I’ve found a LRECL greater than the maximum size Saxon has produced is fine. Similarly a sensible block size works. FILEDATA=TEXT is also needed.

Here’s the SYMNAMES file:

Contents of the SYMNAMES File
RDW,1,4,BI                                                            
Row,%01 
a,%02 

You’ll need an accompanying SYMNOUT DD – for the messages DFSORT (or ICETOOL) produce when the SYMNAMES file is processed.

I’m showing you this first so you can understand the main DFSORT control statements file: Everywhere you see the symbol "Row" in these statements you can interpret it as "%01", whatever that is. Similarly for "a" and "%02". The "RDW" symbol maps the Record Descriptor Word that we need for variable-length record processing. (DFSORT can convert from variable- to fixed-record format but we won’t do that here.)

Now for the control statements:

Contents of the SYSIN File
  OPTION COPY,VLSHRT                                                1
  INCLUDE COND=(1,2,BI,GE,+12)                                      1 
  INREC IFOUTLEN=70,                                                2
    IFTHEN=(WHEN=INIT,                                              3
      PARSE=(%01=(STARTAFT=C'"',ENDBEFR=C'",',FIXLEN=10),           4
             %02=(FIXLEN=8)),                                       5
      BUILD=(1,4,%01,%02)),                                         6
    IFTHEN=(WHEN=INIT,BUILD=(RDW,Row,X,a,SFF,EDIT=(I,IIT)))         7

This is a very simple case of using DFSORT. So, for example, there’s no SORT, no OUTFIL, nor any ICETOOL sophistication. It’s meant to show how you can get the data into a format DFSORT can use. Let me explain how it works:

  1. VLSHRT and the INCLUDE statement will, between them, remove the blank lines Saxon created.
  2. IFOUTLEN sets the output record length (from INREC) to 70 bytes.
  3. This WHEN=INIT parses the input (CSV) data.
  4. The %01 field is filled from after the first " and before the second (with comma) ". It becomes a fixed character field of length 10 bytes.
  5. The %02 field is filled with the remainder of the data in the record – for a length of 8 bytes.
  6. We write out the RDW and both parsed fields.
  7. This WHEN=INIT is used to produce the report lines. We print the %01 field ("Row"), a space, and the %02 numeric field ("a"). For the numeric field ("a") we parse the characters to extract the numeric value (with SFF) and then immediately reformat it (with EDIT=(I,IIT) ) to insert commas.

And here’s the output:

The Resultant Output
One            1                                                      
Two           12
Three        903

Of course we needn’t have just printed the data, as I’ve indicated. With a more interesting data set you could do a lot more.

The use of symbols ("Row" and "a") was largely gratuitous here. It just shows you can use them. If you’re a regular DFSORT or ICETOOL user you’ll know their value.

If you were to strip this down to the bare essentials the first WHEN=INIT does most of the work – parsing the data into fixed positions. (The one really useful thing the second WHEN=INIT does is to convert the numeric field into a packed decimal number.)

So, over these three posts I’ve shown how you can use XSLT to half tame XML data and DFSORT to complete the taming. I have a couple of other things I want to talk about in relation to this. But those belong in a separate post.

XML, XSLT and DFSORT, Part One – Creating A Flat File With XSLT

(Originally posted 2011-05-14.)

This is the second part of a (currently) three-part series on processing XML data with DFSORT, given a little help from standard XML processing tools. The first part – which you should read before reading on – is here.

To recap, getting XML data into DFSORT is a two stage process:

  1. Flatten the XML data so that it consists of records with fields in sensible places.
  2. Process this flattened data with DFSORT / ICETOOL or something else, like REXX.

This post covers the first part of this. You’ll see how you can transform the XML file below into a Comma-Separated Variable (CSV) file.

Here’s the source XML, complete with a few quirks:

XML File To Be Processed
<?xml version="1.0"?>                                                                           
<mydoc>
  <greeting level="h1">
    Hello World!
  </greeting>
  <stuff>
    <item a="1">        1
      <row>One</row>
    </item>
    <item               2
      a="12">
      <row>Two</row>
    </item>
    <item a="903">      3
      <row>
      Three
      </row>
    </item>
  </stuff>
</mydoc>

Here’s the resulting flat file:

Resulting Flat File For Processing With DFSORT / ICETOOL
               
    "One",1                                                                  
    "Two",12                                                                                    
    "Three",903                                                                           
               

I’m assuming you can read XML reasonably well. In this example we have three "item" elements as children of a "stuff" element. The "stuff" element is a child of the "mydoc" element. The "mydoc" element also contains a "greeting" element. Each "item" element has a single "row" child element and an "a" attribute.

To produce the output we need to find the "item" elements and pick up the "row" child element and the "a" attribute value. We write one record for each "item" element. (We ignore the "greeting" element entirely.)

You may notice some white space around the output: A leading blank line and a trailing one, as well as four spaces at the beginning of each output record. I’ve not found a way for getting rid of those and the DFSORT program (described in the next part of this series) will have to strip them off.

I’ve deliberately formatted each "item" element slightly differently:

  1. The "a" attribute is on the same line as the "item" tag, and the "row" element fits entirely on one line.
  2. The "a" attribute is on the next line, and the "row" element is on one line.
  3. The "a" attribute is as in 1 but the "row" element text is split across three lines.

The point is that XML is so flexible in its layout you’re better off relying on a supplied parser than writing your own. It’s true that there are good parsers that don’t do XSLT transformations. And obviously the z/OS System XML one is very nice, particularly with its ability to use specialty engines. As I said in my previous post, XML parsing is computationally expensive.

Why not write your own code that calls the z/OS System XML parser? That’s certainly an option – and indeed you might find the transformations you want to do can’t (or shouldn’t) be done with XSLT. Here the similarity to DFSORT is quite strong: Both provide ways to use built-in functions to transform data – neither of which require a formal programming language (in XML’s case perhaps PHP, java or C++ and DFSORT’s case perhaps Assembler, COBOL or PL/I).

In this example you scarcely need to write your own program. (Handling item 3, as I’ll describe later, is the one case where a program might be better.).

Here’s the XSLT stylesheet that produces the required output:

XSLT Stylesheet
<?xmlΒ version="1.0"?>
<xsl:stylesheetΒ version="2.0"                               1                                    
Β Β xmlns&colon;xsl="http://www.w3.org/1999/XSL/Transform">
Β 
Β Β <xsl:outputΒ method="text"Β encoding="IBM-1047"/>           2
Β 
Β Β <xsl:templateΒ match="/">
Β Β Β Β <xsl:apply-templatesΒ select="mydoc/stuff"/>             3
Β Β </xsl:template>
Β 
Β Β <xsl:templateΒ match="item">                               4
Β Β Β Β <xsl:text>"</xsl:text>                                  5
Β Β Β Β <xsl:value-ofΒ select="normalize-space(row)"/>           6
Β Β Β Β <xsl:text>",</xsl:text>                                 7
Β Β Β Β <xsl:value-ofΒ select="@a"/>                             8
Β Β </xsl:template>
Β 
</xsl:stylesheet>

This is a fairly simple stylesheet. Here’s how it works (and the numbered lines above correspond to the numbering below:

  1. Here we declare the level of the XSLT language to be 2.0. In fact there’s nothing about this stylesheet that requires that language level.
  2. Here we say we’re creating a text file as output and that it will be EBCDIC (IBM-1047).
  3. Here we search for the "stuff" element within the "mydoc" element – using the XPath language. In fact the only "stuff" element we’ll match with is the one at the top of the XML node tree – because it’s preceded by a "/". For each matched "stuff" element we apply the template below.
  4. This template matches all "item" elements within the "stuff" element.
  5. Here text starts to be written out for the record. In this case the leading quote around the first piece of data.
  6. Here the first of piece of data is written out – the text value of the "row" element. We’ll come back to the normalize-space() function in a minute.
  7. Here a trailing quote and a comma are written out.
  8. Here the value of the "a" attribute is written out. It needs no adjustment (in this example).

Because item 3’s "row" value was split across several lines the normalize-space() function is used to take out leading white space. It has the unfortunate side-effect of replacing multiple white space characters in the text with a single space so it’s not brilliant. You could write a fairly simple but recursive piece of XSLT to do the job properly – but it’s beyond the scope of this post. In fact this might be the thing that makes you abandon XSLT and call the XML parser from a program.

If you want to get into XSLT I can recommend Doug Tidwell’s XSLT, Second Edition Mastering XML Transformations book. It’s what I’ve used – with some additional research on the web (which didn’t yield much additional insight).

I used the Saxon B (free) parser as it’s the only one I can get my hands on that does XSLT 2.0. It’s a java jar. You could use others, of course.

Invoking from the OMVS I found a 64MB heap specification was enough (running in a 128MB region). For more complex transformations I can see a larger heap might be needed. (In fact I didn’t check how much garbage collection, if any, the JVM did. It just ran.) :-)

(If you specify version="1.0" for the stylesheet Saxon will issue a message informing you you’re running a 1.0 stylesheet through a 2.0 processor. This has caused no problems whatsoever for me.)

Originally I downloaded Saxon to my Linux laptop and used it with an ASCII stylesheet and XML data. Transferring to z/OS was straightforward. This approach may work for you, if you’re setting out to learn XSLT.

Learning and working with XSLT continues to be a journey of discovery. If I’m missing some tricks that you spot feel free to let me know. The next post in this series will be about the DFSORT counterpart.

XML, XSLT and DFSORT, Part Zero – Overview

(Originally posted 2011-05-11.)

In the distant past I’ve written about using DFSORT to parse XML. This post (and two follow-on posts) will describe an experiment to make such processing much more robust.

In this post I’ll talk about what the problem I’m trying to solve is. And why. And a brief outline of my solution.

About XML

This isn’t meant to be the most detailed description of XML, nor a complete list of where it’s used. I just want you to know (if you didn’t already) why I think XML processing is something to pay attention to.

Increasingly applications are producing and consuming XML. (They’re also producing and consuming other new data styles, such as JSON.) I divide this usage into two categories:

  • Configuration data (generally small files).
  • Business data (often very large files).

XML has many advantages as a data format, including robustness, standardisation and an increasing degree of inter-enterprise adoption. It also has useful attributes like the ability to validate a file against a strict grammar and also transformability.

XML is, however, expensive to parse. And when I talk of transformability the tools to transform XML are still quite rudimentary – you often have to write your own program to do it.

(This being an IBM-hosted blog you might expect me to talk about Websphere Transformation Extender (WTX). I shan’t, except to say it has very nice tooling. Similarly, you might expect me to talk about the Extensible Stylesheet Language for Transformations (XSLT) – as a standard for transformations. You’re in luck with XSLT – but that will have to wait. I’d like to talk about IBM’s z/OS XML Toolkit (which includes an XSLT processor) but that will have to wait. And as for DataPower, it’ll be a while before I talk about it, also.)

Those of you familiar with IBM mainframe technology will be aware of z/OS System XML and perhaps the z/OS XML Toolkit. You’re probably aware of the ability to offload XML parsing to a zAAP (or zIIP if zAAP-on-zIIP is in play). I think our story’s pretty good with these.

So IBM thinks XML’s important, and so do lots of installations. It’s important that mainframe people know what they can do, too.

The Problem I’m Trying To Solve

I don’t feel it necessary to describe what DFSORT can do in this post. Suffice it to say it can do lots of what I call "slice and dice" with data. So long as that data is record-oriented. (And it’s even better if you include ICETOOL.)

So why don’t we just process XML with DFSORT?

(Let’s disregard publishing XML with DFSORT as that’s very easy to do.)

Traditionally DFSORT has done really well when records are neatly divided into fixed-position (and length) fields. Over recent years it’s got better and better at handling cases where the layout of each record is variable. For example, it can parse Comma Separated Value (CSV) files just fine – with PARSE.

But XML is so much more variable. For example, two partners could each send you a file, created by their own programming or tools. They’d be semantically equivalent but the data would be differently formatted (and still be valid according to the same XML Schema). And the differences wouldn’t just be the fields being at different offsets, or in a different order in the same record: One format might have an element all on one line whereas the other might spread it across three lines.

So any DFSORT application attempting to process XML would be vulnerable to this variability. In the past, when I’ve written of DFSORT processing XML I think I’ve said that you need stable XML to work with. I think that’s still right.

So is that it? Well, no it isn’t: I still think it’s possible to take advantage of DFSORT’s power, even with XML data to process. Read on…

XSLT

XSLT (standing for Extensible Stylesheet Language for Transformations) is a standards-based way of transforming XML – to (different) XML, HTML or even plain text. And by "(different) XML" I also mean things like SVG vector graphics.

With XSLT you define a transformation using another piece of XML – a stylesheet (or XSL file). Whether you author this by hand (my current state) or use tooling to generate one is up to you. Using a program you use the XSL file to transform your XML to whatever you want.

There are lots of XSLT programs. I’ve used Apache Xalan (which is tightly-coupled to the IBM ones on z/OS), Saxon, the capabilities built in to Firefox (and other browsers), PHP’s one – to name just a few. Of these only Saxon can do XSLT 2.0 at present. (The others all do XSLT 1.0, often with extension capabilities.)

For my work, written up in these posts, I used the free variant of Saxon – because it does 2.0. Nothing in these posts, however, requires 2.0. I want 2.0 just so I can learn 2.0. One day maybe it’ll catch on and then I’ll be in good shape. Learning 2.0 isn’t incompatible with learning 1.0 but it might leave you frustrated. πŸ™‚

The important piece in all this is that XSLT can be used to take arbitrary XML and flatten it – into records with fields in vaguely sensible places. In EBCDIC.

Putting It Together

So far I’ve talked about two distinct components: DFSORT / ICETOOL and XSLT. I’ve said it’d be nice to be able to process XML-originated data using DFSORT, robustly. So here’s how it can be done:

  1. Use XSLT to create a flat file (in HFS or zFS) with the data flattened into sensible records with well-delimited fields. (In the example, in the next post in this series, I’ll use CSV as the intermediate file layout.)
  2. Use DFSORT’s parsing capabilities to read the intermediate file and then do DFSORT’s normal things with it. (This will be the third post in the series.)

Conceptually simple but a little fiddly in the details. In the next two posts I’ll clothe the idea with some of those details.

Over the past few days, while preparing to write this post, I’ve done some experimenting – including creating a full working example. There are lots of "wrinkles" on this idea, including other ways of doing pieces of it. Perhaps you’ve thought of a few. If so do let us know.

Vienna Conference – A Trip Report

(Originally posted 2011–05–09.)

I think people know better than to ask me for a trip report to a conference I’ve attended. They’ll get what I think is important – and their priorities are probably different. So here is that trip report anyway… πŸ™‚

You’ll probably have gathered by now I’m for a “for the journey” person than a “for the destination” one. But I won’t bore you with the minor inconveniences on both ends of the trip – because I personally try to forget the (often long and tedious) journey when I get to the destination (or home again). I’d rather focus on where my travels took me.

I will admit to sampling hostelries – with good friends. I also was very pleased to be in the company of friends – both old and new. Personally, I think the social aspect of a conference is almost as important as the sessions. And, of course, some really useful conversations were had – with IBMers, business partners, vendors and customers. I can’t really summarise these – for the usual obvious reasons.

My four sessions mainly went well. The topics are summarised here. But here are my perceptions:

  • “Parallel Sysplex Performance Topics” went well, I think. Mainly because I talked about the subset of items I really wanted to talk about. Most notably “Structure Execution Time” and “Structure Duplexing Performance”. (And I had a very good question on how the non-CPU element of request time relates to distance and technology.)

  • I think “Much Ado About CPU” has become disorganised. It needs refocusing. Particularly as I expect the CPU picture to continue to evolve over time. And so this one has to survive in some form.

  • “Memory Matters” was done while too tired. I also think it contains too much baggage from DB2 Version 8 (even though many customers are still on 8). I also think the “Coupling Facility Memory” section doesn’t really add much.

  • “DB2 Data Sharing Performance For Beginners” turns out not to be a “for beginners” presentation, really. If I’m introspective about it I thought when I wrote it it would help explain the major themes but I couldn’t pretend be as knowledgeable as the true greats of Data Sharing. For example, those that write “DB2 Performance Topics” Redbooks. So I should skip the “for beginners” part of the title and rework it to make it as good a presentation as I can for those who already have some knowledge. The stuff needs saying but I need to say it better.

But, I think in the above I’m being harsh on myself. I got good evaluations on all four. Maybe the audience is very kind. πŸ™‚

I took notes using the Writepad handwriting application on iPad (into Evernote so I can read them and edit them everywhere). Writepad does a very good job but I still wish I’d brought the keyboard along: I found the mechanics of taking notes diminished my ability to listen. I’d pull out the following presentations as ones I got a lot out of. (Others will have their own favourites.)

  • Susann Thomas (a team-mate from the 2009 Batch Modernisation residency) did a very nice job on introducing XML for System z. So much so I’m convinced I need to understand the XML story better. (You may have seen on Twitter my attempts to do stuff.)
  • Harald Bender’s XML and RMF presentation makes me think a practical example of XML to play with is that produced by RMF.
  • Marna Walle did a nice job of her z/OS R.13 Preview presentation. (Which reminds me I must write on in-stream SYSIN in a PROC soon.)
  • George Ng (who apparently reads this blog ! πŸ™‚ ) presented on Infiniband Coupling Facility links. I note all RMF knows about Infiniband links is the channel path acronym “CIB”. It can’t distinguish between e.g 1x SDR and 12x DDR, for example. You can imagine I’d “have views” on that sort of thing. :-)”)
  • Christian Daser explained rather well, I thought, the tricky DB2 V10 Bitemporal support, as well as a few other pieces of DB2 Application componentry in 10.
  • Peter Enrico will certainly have opened some eyes to the value of SMF 113 CPU Measurement Facility instrumentation. I’ve been familiar with this for a long time – certainly from before we announced it. I would write about it if I didn’t feel Peter (and John Burg) hadn’t already done so as well as I could have – if not better.
  • Mike Buzzetti gave a very good introduction to Cloud on System z, particularly about TSAM for provisioning.
  • I’ve tried to run with Jeff Berger’s foils before now. I’m so glad I don’t have to anymore: He does so much better a job of it than I do. πŸ™‚ The topic I saw him present this time was on DB2 V10 Performance. I’m eagerly awaiting the Redbook, of course.
  • And last but not least Bob Rogers’ “What You Do When You’re a z196 CPU”. I’m very glad he keeps updating it for each generation of processors. It’s one where you really do need to know what happened before so I’m pleased he’s kept in z9 and z10 stuff.

Of course I don’t know whether you have access to the proceedings. If you do I recommend you pull down some of the above sets of slides. If not maybe you’ll see them at some other conference or user group.

After a week of this I’ll admit to coming home very tired. (In fact I think everyone felt that way by Thursday morning.) But it was a great week for me. And thanks to everyone who made it so good for me.

And if you didn’t get to Vienna I hope you do get to some System z conferences: They’ve a very good use of money and time.

Batch Architecture, Part Three

(Originally posted 2011-05-04.)

Up until now I haven’t talked much about DB2, except perhaps to note it’s a little different. But what is a DB2 Batch job anyway? It’s important to note a DB2 job ISN’T necessarily exclusively DB2 – although some are. It’s just a job that has some DB2 in it.

The reason for writing a separate post, apart from breaking things up a little, is because batch jobs with DB2 in them present particular challenges. But also additional opportunities. In general these jobs can be treated like others but with extra considerations.

The main challenge is determining which data the job accesses – and how it accesses it. Let’s break this up into two stages:

  1. Identifying which DB2 plans and packages are accessed by which job / step.
  2. Identifying which DB2 tables and other objects are used by these plans and packages. And perhaps how.

Identifying DB2 Plans and Packages

This piece is relatively straightforward: DB2 Accounting Trace -with trace classes 7 and 8 enabled – will give you the packages used. You need to associate the Accounting Trace (SMF 101) record with its job / step.

For most DB2 attachment types the Correlation ID is the same as the job name. (Identifying the step name and number is a matter of timestamp comparison with the SMF30 records – which my code learned to do long ago.)

For IMS it’s more complicated, with the Correlation ID being the PSB name.

(A byproduct of this step might be discovering which jobs use a particular DB2 Collection or Plan name. Sometimes these are closely related to the application itself.)

Identifying Used Objects

This piece is much harder, particularly for Dynamic SQL. Fortunately most DB2 batch uses Static SQL. Even so it’s still tough: If you have the package names you can use the DB2 Package Dependency table in the Catalog to figure out which tables and views the package uses. At least in principle: There’s no guarantee these dependencies will get exercised – as there’s no guarantee the statements using them will ever get executed.

Another problem with this is figuring out whether the access is read-only or for-update.

To totally figure out which statements are executed (and which objects they update and read) would require much deeper analysis – probably involving Performance Trace and extracting SQL statement text from the Catalog.

Conclusion

So this is very different from the non-DB2 case. But at least we can glean what data a DB2 batch job OUGHT to be interested in. And, by aggregation, it’s not hard to work out what data an entire batch application uses.

In this post I wanted to show how DB2 complicates things but that it’s not hopeless. In fact there’s a substantial silver lining to the cloud: Without examining the (possibly missing) source code you can look inside the job at the embedded SQL, if you’re prepared to extract them from the DB2 Catalog.

You’ll notice I’ve said very little in this set of posts about Performance. This is deliberate: Although much of the instrumentation I’ve described is primarily used for Performance these posts have been about Architecture. Which is, I think, a different perspective.

I expect I’ll return to this theme at some point. For now I’ll just note it’s been fun thinking about familiar stuff in a slightly different way.

By the way this post was written using the remarkably accurate WritePad app on the iPad. It’s grown better at recognising my scrawl in the few hours I’ve used it – or perhaps it’s me that’s getting trained. πŸ™‚

I Know What You Did Last Summer

(Originally posted 2011-04-26.)

This is literally a sketchy outline for a new presentation I want to build. The working title is indeed "I Know What You Did Last Summer".
There’s clearly not much structure to this. But the basic outline idea is there: What can an installation glean without too much effort?

Let the graphology begin. πŸ™‚

Batch Architecture, Part Two

(Originally posted 2011–04–25.)

I concluded Batch Architecture – Part One with a brief mention of inter-relationships and data. I’d like to expand on that on this part.

Often the inter-relationships between applications are data driven – which is why I’m linking the two in this post (and in my thinking). But let’s think about the inter-relationships that matter. There are four levels:

  1. Between applications.
  2. Between jobs.
  3. Between steps in a job.
  4. Between phases in a step.

The first three are well understood, I think. The fourth is something I explored last year. Before I talk about it let me talk about “LOADS” – which I mentioned in Memories of Hiperbatch.

(And a minor note on terminology: Yes I KNOW that OPEN and CLOSE are macros. I don’t intend to use the capitalisation here – because the act of opening and closing a data set is meaningful, too (and less grating to read). Forgive me if this “sloppiness” offends.) :-)")


Life Of A Data Set (LOADS)

I won’t claim to have invented this technique. (As I said in “Memories of Hiperbatch” I declined an offer to write up a patent application because I knew I hadn’t originated it.) But I do advocate its use quite a bit. Here’s an (oft-used) example:

If you have a single-step job that writes a sequential data set and another that reads it (both from start to finish) there’s a characteristic data set “signature”: Two opens, one after the other, one for update, one for read. If you discern this pattern you might think “BatchPipes/MVS”. (Depending on other factors you might think other things – such as VIO.)

So this is a powerful technique.


LOADS Of Dependencies :-)")

In 1993 we wrote code to list the life of each data set a job opened and closed. Not long after that we got tired of figuring out dependencies by hand from LOADS. :-)") So we fixed it:

At its simplest a writer followed by a writer indicates a (“WW”) dependency. A writer followed by a reader indicates a (“WR”) dependency, also. And so on.

Pragmatically some of these dependencies aren’t real, or at least it isn’t as simple as this sounds. For example:

  • This says nothing about PDS members.

  • GDGs are a little different.

  • A writer one morning and a reader the same evening might not be marked as a dependency in the batch scheduler (though it probably ought to be). To at least alert the analyst (mainly me these days) to this sort of thing the code pumps out the time lag between the upstream close and downstream opens. (This is an enhancement I made, together with some more “eyecatcher” things with timestamps last year.)

  • What’s the key here? Do we include volser?

But you can see there’s lots of merit in the technique, even with these wrinkles.


Step Phases

As I said before application-level, job-level (in some ways the same thing) and step-level dependencies are things we’ve known about for a long time. Also we’ve know about DFSORT (and other sort) phases for a long time, too: Input, Intermediate-Merge and Output phases. These should be familiar, although people tend to forget about the possibility of an intermediate merge phase – because it should only apply to large sorts.

So, if sorts have phases, what about other steps? Last year I enhanced the code to create Gantt charts for data set opens and closes within a step. In many cases jobs became no more interesting because of it. But in a number of cases fine structure appeared: Non-sort steps demonstrably had phases. In one example a step that read a pair of data sets in parallel wrote to a succession of output data sets. I could see this from the open and close timestamps of the output data sets. (Without looking at the source code I couldn’t be sure but maybe there’s some mileage in dissecting this step.)

It’s in my code: If it applies to your jobs I’ll be sure to tell you about it.


An Application And Its Data

Apart from the small matter of scale figuring out which data an application uses is the same problem as figuring out which data a job uses.

I think I’ll talk about DB2 in a later post, as this one has already become lengthy.

As you probably know there is lots of instrumentation on data sets in SMF. Without going into a lot of repetitive description:

  • You can get information about disk data sets from SMF 42 Subtype 6.
  • You can get information about VSAM data sets from SMF 62 (open) and SMF 64 (close).
  • For non-VSAM it’s SMF 14 (for read) and 15 (for update).

There are a number of lines of enquiry you might like to pursue, including:

  • Working out which data sets contribute most to the application’s processing time.

    Here you’d use SMF 42 and something like I/O number or (more usefully) I/O number times Response Time.

  • Figuring out which data sets are strongly related to this application and no other.

    In this case SMF 14, 15, 62 and 64 are needed. (You don’t need both 62 and 64 for the same data set.)

None of the above applies to DB2: You don’t get 14, 15, 62 or 64 for DB2 data (despite DB2 using Linear Data Sets, a form of VSAM). But there is useful work you can do on DB2 data classification. And that is the subject of the next post in this series.

DFSORT – Now With Extra Magical Ingredients

(Originally posted 2011–04–21.)

Thanks to Scott Drummond for reminding me of last Autumn’s DFSORT Function PTFs – UK90025 and UK90026. They’re mentioned in the preview for z/OS Release 13 so now is not such a bad time to be talking about them. So let me pick out a few highlights:



Translation Between ASCII And EBCDIC, And To And From Hex and Binary

For a long time DFSORT has been able to translate to upper case (TRAN=LTOU), to lower case (TRAN=UTOL) and using a table (TRAN=ALTSEQ).

Now you can convert from ASCII to EBCDIC (TRAN=ATOE) and back (TRAN=ETOA). Translation is performed using TCP/IP’s hardcoded translation table.

Other “utility” translations are added: BIT, UNBIT, HEX and UNHEX. For example TRAN=HEX would translate X’C1F1′ to C’C1F1′ and TRAN=UNBIT would
translate C’1100000111110001′ to X’C1F1′.



Date Field Arithmetic

DFSORT already had some nice functions for handling dates and times. But here are some new things. This isn’t an exhaustive list:

  • You can add years to a date field – with ADDYEARS.
  • You can subtract months from a date field – with SUBMONS.
  • You can calculate the difference between two dates – with DATEDIFF.
  • You can calculate the next Tuesday for a date field – with NEXTDTUE.
  • You can calculate the previous Wednesday for a date field – with PREVDWED.
  • You can calculate the last day of the quarter – with LASTDAYQ.


JCL Symbols In Control Statements

You can now construct Symbols incorporating JCL PROC or SET symbols. These can be used in DFSORT and ICETOOL control statements, just like other symbols. You specify this by coding JPn“&MYSYM” in the PARM parameter of the EXEC statement. (In fact there need be no JCL symbol in this so you could pass in other strings this way and an expected use is for JPn to contain a mixture of JCL Symbols and fixed text.) n can be any one of 0 through to 9.

This support is in addition to the ability to use System Symbols (introduced with UK90013).


Microseconds In Timestamps

You can use the new DATE5 keyword to create a timestamp constant at run-time in the form ‘yyyy-mm-dd-h.mm.ss.nnnnnn’. DB2 folks might recognise this as the timestamp format for DB2 Unload and DSNTIAUL. You can use this for things like comparisons.


Chunking And Stitching Together Records

You can use the new ICETOOL RESIZE operator to:

  • Split records into fixed-sized output records. For example, take a RECFM=FB, LRECL=500 file and create a RECFM=FB, LRECL=100 file – creating 5 new output records from each input record.
  • Join together fixed-sized input records. For example, take a RECFM=FB, LRECL=100 file and create a RECFM=FB, LRECL=500 file – in effect reversing the above by joining 5 input records together to make an output record.

In each case you can see there could be problems with partial output records. DFSORT “does the right thing” using blanks.


Begin Group When Key Changes

I mentioned WHEN=GROUP here, in particular BEGIN=. With BEGIN= you get a new group when the condition you specify is satisfied. Now, with KEYBEGIN= you get a new group when the value in a particular field changes. For example:



SORT FIELDS=(1,12,CH,A,13,8,CH,D)
OUTREC IFTHEN=(WHEN=GROUP,KEYBEGIN=(1,12),
PUSH=(13:13,8,31:ID=3))






sticks a group number (3 characters wide) on the end of each record. The group number is incremented when there’s a new value in the 12-byte field that begins in position 1.


There are lots of other (in my opinion) smaller enhancements in this PTF.

If you want to know whether the appropriate PTF is on look for the following message in a DFSORT run:



ICE201I H RECORD TYPE …



If you see an “H” then you’re all set.

And you can read about these enhancements in more detail in User Guide for DFSORT PTFs UK90025 and UK90026 (SORTUGPH).