Thursday, April 17, 2014

Tracking the clickahertz and jiggarams - Part 1

I assembled this list some time ago with every intention to turn it into actual Collections and Reports in PerfMon on all my IIS boxes. This has yet to happen, mainly because there were a lot of ongoing discussions about systems monitoring, Epic's SystemPulse product, BMHCC's enterprise SCOM implementation, and others.  And also because my mind - and, thus, sense of time and prioritization - is a hot mess.

Cobbled together from various source, this list represents a best first attempt at a universally applicable list of key Windows metrics regardless of a server's stated purpose.


My plan for "Part 2" involves a similar collection of metrics specific to IIS and ASP.NET applications. Those waters are considerably harder to navigate for me as a non-developer, but they're no less critical to answering questions about concurrent client connections, number of unique requests to an application, tracking memory leaks or other weirdness in worker processes, etc. I'll post that once I have it.  

Without further ado:

Memory

Counter Description Use Notes
Memory\Available Mbytes Available system memory in megabytes (MB) <10% considered low

<5% considered critically low

10MB negative delta per hour indicates likely memory leak
Memory\Committed Amount of committed virtual memory. Also called the "commit charge" in TaskManager
Memory\System Cache Resident Bytes Amount of memory consumed by the system file cache. Shows as "Metafile" in a memory explorer.

In 64-bit systems, memory addressing allows system file cache to consume almost all physical RAM if unchecked.
Memory\Pages Input/sec Rate of total number of pages read from disk to resolve hard page faults >10 considered high

Compare to Memory\Page Reads/sec to determine number of pages read into memory for each read operation.

Will be => Page Reads/sec, and large delta between them might indicate need for more RAM or smaller disk cache
Memory\Pages/sec Rate of total number of pages read from and written to disk to resolve hard page faults >1000 considered moderate, as memory may be getting low

>2000 considered critical, as system likely experiencing delays due to high reliance on slow disk resources for paging

Sum of Pages Input/sec and Pages Output/sec
Memory\Page Faults Total number of hard and soft page faults Pages/sec can be calculated as a % of Page Faults to determine what % of total faults are hard faults
Process(*)\Handle Count Total count of concurrent handles in use by the specified process. As this number constantly fluctuates, the delta between high and low values is most important.

Max - Min = !> 1000

Consistent large number or aggressive upward trend of handle count commonly causes memory leaks

  
Processor
Counter Description Use Notes
System\Processor Queue Length Total number of queued threads waiting to be processed for all processors. =>10 considered high for multi-processor system

PQL/n, where n = number of logical processors, gives per-core queue length.

If % Processor Time is high (=>90%) and per-core PQL is =>2, there is a performance bottleneck.

It is not uncommon to experience low % Processor Time and a per-core PQL => 2 depending on efficiency of the requesting application's threading logic
Processor\%Processor Time Percentage of time the specified CPU is executing non-idle threads. >75% considered moderate and should be closely monitored

>90% considered high, may begin to cause delays in performance

>95%-100% considered critically high and will cause major delays in performance

This value should be tracked per logical processor.

If % Processor Time is high (=>75%) while disk and network utilization is low, consider upgrading or adding processors.
Processor\%Interrupt Time Percentage of time the specified CPU is receiving and servicing hardware interrupts from network adapters, hard disks, and other system hardware Interrupt rates 30%-50% or higher may indicate a driver or hardware problem


Disk I/O

Counter Description Use Notes
PhysicalDisk\%Disk Time Percentage of time the selected disk spends servicing read or write requests. If this value is high relative to nominal CPU and network utilization figures, it is likely disk performance is a problem.
PhysicalDisk\Disk Writes/sec Average number of disk writes per second Used in conjunction with Disk Reads/sec, general indicator of disk I/O activity
LogicalDisk(*)\Avg Disk sec/Writes Average time in seconds a specified disk takes to process a read request >15ms considered slow and worth close evaluation

>25ms considered very slow and likely to negatively impact system performance
PhysicalDisk\Avg Write Queue Length Average number of write requests waiting to be processed Used in conjunction with Avg Read Queue Length, this gives an idea of disk access latency.

AWQL/n <= 4, where n is the number of disks in RAID.
PhysicalDisk\Disk Reads/sec Average number of disk reads per second Used in conjunction with Disk Writes/sec, general indicator of disk I/O activity
LogicalDisk(*)\Avg Disk sec/Reads Average time in seconds a specified disk takes to process a write request >15ms considered slow and worth close evaluation

>25ms considered very slow and likely to negatively impact system performance
PhysicalDisk\Avg Read Queue Length Average number of read requests waiting to be processed Used in conjunction with Avg Write Queue Length, this gives an idea of disk access latency.

ARQL/n <= 4, where n is the number of disks in use.


Network I/O

Counter Description Use Notes
Network Interface(*)\Total Bytes/sec Measure of total bytes sent and received per second for the specified network adapter If >50% of Current Bandwidth value under typical load, problems during peak times are likely.
Network Interface(*)\Current Bandwidth Estimate of the current bandwidth in bits per second (bps) available to the specified NIC.

Considered "nominal bandwidth" where accurate estimation impossible or where bandwidth doesn't vary
To estimate current NIC utilization, use the following formula:

Nic Utilization = ((Max Bytes Total/Sec * 8) / (Current Bandwith)) * 100

>30% NIC utilization on shared network considered high
Network Interface(*)\Output Queue Length Total number of threads waiting for outbound processing by the specified NIC >1 sustained considered high

>2 sustained considered critically high

Tuesday, April 15, 2014

From (F)ailure to (A-)wesome on SSLLabs

I'm once again shamelessly copy/pasting a new post on here.  I don't feel too ashamed of it, since I did actually write the original post.  It turns out we have MySites at work, and that the Blog feature is enabled!  I'll very likely be posting in both places as the life of this Blog thing draws onward.
 ***

There's been a lot of chatter about the Heartbleed SSL vulnerability in the last couple of weeks, and rightfully so. One place folks seem to love going is over to SSLLabs, since they have a server tester you can run to determine what kind of safety grade – A to F – you get.
At the outset, my tests of the BOC Link and MyChart sites generated giant, terrifyingly red "F" results. This was not due to Heartbleed, thank goodness, since the NetScalers do not use an affected version of OSSL, and none of my web servers use OSSL at all. What failed me instead was another, slightly older vulnerability: SSL renegotiation.
Both BOC Link and MyChart run behind a NetScaler virtual VPX appliance running v10.0.x of the software. Out of the box, NetScalers are configured to allow all SSL renegotiation in all forms, whether initiated from the client connection or the server. A quick check at the console will tell you the current status of the parameter:

> show ssl parameter
Advanced SSL Parameters
-----------------------
SSL quantum size: 8 kB
Max CRL memory size: 256 MB
Strict CA checks: NO
Encryption trigger timeout 100 mS
Send Close-Notify YES
Encryption trigger packet count: 45
Deny SSL Renegotiation NO
Subject/Issuer Name Insertion Format: Unicode
OCSP cache size: 10 MB
Push flag: 0x0 (Auto)
Strict Host Header check for SNI enabled SSL sessions: NO
PUSH encryption trigger timeout: 1 ms
Global undef action for control policies: CLIENTAUTH
Global Undef action for data policies: NOOP 


Citrix has a pretty handle article on what exactly the –denySSLReneg parameter is, what its options are, and how to change it. See it here.
Here's the command:

> set ssl parameter -denySSLReneg NONSECURE
Done 

By setting the Deny SSL Renegotiation option to NONSECURE, I've corrected the renegotiation vulnerability without (hopefully) creating any compatibility issues for our Link and MyChart users. This setting appears to be global, so affecting this change raised the scores of both sites from "F" to "A-" (RC4 ciphers, indeed!) simultaneously.

> show ssl parameter
Advanced SSL Parameters
-----------------------
SSL quantum size: 8 kB
Max CRL memory size: 256 MB
Strict CA checks: NO
Encryption trigger timeout 100 mS
Send Close-Notify YES
Encryption trigger packet count: 45
Deny SSL Renegotiation NONSECURE
Subject/Issuer Name Insertion Format: Unicode
OCSP cache size: 10 MB
Push flag: 0x0 (Auto)
Strict Host Header check for SNI enabled SSL sessions: NO
PUSH encryption trigger timeout: 1 ms
Global undef action for control policies: CLIENTAUTH
Global Undef action for data policies: NOOP

Monday, January 6, 2014

I hate dirty Application Logs - PerfMon counters and IIS Advanced Logging

I posted this one on Epic's UserWeb entity portal and have in my guilt for not posting in a while blatantly copied and pasted it here.  Methods aside, it's good info, especially if you're OCD about what shows up in your server's error logs like I apparently am.  Source link at the bottom.

Enjoy!

-------------
I have been configuring IIS Advanced Logging on all the IIS servers I've been building ahead of our go-live in January. It works swimmingly and solves a couple of problems I'd always hated about standard IIS logging:
1. The logging happens in real-time, instead of on a 3 minute delay.
2. You can add custom fields, like "Client-IP," that work a lot more smoothly with ADC's and load balancers that might otherwise mask information about a logged session.
3. You can include basic performance counters in your logs, like the W3WP CPU and memory utilization.

That #3 is why I'm posting here tonight. Even though those fields were disabled in my default log definition, I'd still get the following error for each metric in my Windows Application log:
____
Log Name: Application
Source: IIS Advanced Logging Module
Date: 12/17/2013 10:22:29 AM
Event ID: 1008
Task Category: None
Level: Error
Keywords: Classic
User: N/A
Computer: ECEPPRINTC01.ad.bmhcc.org
Description:
Failed to initialize performance counter \Process(w3wp)\Private Bytes. Data for this performance counter data will not be recorded until the counter is available. PdhCollectQueryData: 0x0X800007D5.
Event Xml:
<Event xmlns="http://schemas.microsoft.com/win/2004/08/events/event">;
<System>
<Provider Name="IIS Advanced Logging Module" />
<EventID Qualifiers="0">1008</EventID>
<Level>2</Level>
<Task>0</Task>
<Keywords>0x80000000000000</Keywords>
<TimeCreated SystemTime="2013-12-17T16:22:29.000000000Z" />
<EventRecordID>2513</EventRecordID>
<Channel>Application</Channel>
<Computer>MYSERVERNAME</Computer>
<Security />
</System>
<EventData>
<Data>\Process(w3wp)\Private Bytes</Data>
<Data>0X800007D5</Data>
</EventData>
</Event>
____
The short of it: this error showed up in the logs of those servers whose application pools I'd configured to use ApplicationPoolIdentity to authenticate instead of the old standby NetworkService. It occurs because ApplicationPoolIdentity has no rights to Performance Monitor, and so no access to log using Performance Monitor counters.

The fix is to add your application pool's identity ("IIS APPPOOL\APPPOOLNAME") to the built-in Performance Monitor Users group. Doing so eliminates the errors in Windows Application log, and makes the metrics actually show up correctly in the log (instead of "-" like they were originally).

Here's where I found the info:
http://blogs.microsoft.co.il/idof/2013/08/20/fixing-iis-advanced-logging-performance-counters-errors/

Sunday, November 3, 2013

The Tits: IIS Advanced Logging Module

The first time I thought "Wow, NetScalers sure fill my IIS server logs with a bunch of shit data" was about twenty minutes after my first encounter with the technology.  I had checked every setting I could think of on the IIS Logging feature and found nothing of use.  It's an all or nothing proposition: enabled, or off completely.

Tonight, the clouds parted, and the moon shone brightly.

I came in tonight to try to get some things ready for some server build audits we have coming up.  My two main tasks were as follows:
  1. Get three of the core web application servers built on the NetScaler (that I finally have access to) and load balanced in a basic capacity.
  2. Get the current version of the application deployed to those same three servers.
The first task I accomplished pretty handily, despite my apprehension logging into a production NetScaler for the first time.  I've been through the training, and I've been reading through their eDocs, of course, but I went very slowly for the first few minutes. 

I had not long completed #1 and successfully tested it when my thoughts again returned to my servers' IIS logs.  They were filled with 0-byte HEAD requests from the NetScaler.  By default, the NetScaler's built-in http monitor executes a HEAD method request against the target system every five seconds.  I knew this from earlier study and my general understanding of monitoring systems.  What I had not considered before the last month or so was what this method of monitoring looks like on the receiving end.

I turned to the Intergoogle for aid. There had to be an answer to this problem.  The legions of IIS administrators that I still think might be out there somewhere surely could not have, all this time, merely tolerated this situation.

Right?

Enter, the solution

There is, indeed, a solution to this problem, and it is called IIS Advanced Logging.  I've only just begun to explore its capabilities; however,I determined almost immediately that I must have this module on my servers.

All of them.

The reason is simple: it lets me filter out unwanted log traffic.  Further, should I determine a sound need to (which I'm still researching), I can separate out different log data into separate files based upon what I want to actually see.

If this module did nothing else of value, this one feature would justify its existence.  From the reading I've already done, though, this thing is all kinds of flexible.  There are some limitations, but at least for now they do not affect what I'm doing.

I have installed, enabled, configured, and tested AdvL on all three of tonight's target servers, and it's running like a champ.  I did find it odd that you have to manually disable standard IIS Logging, else it logs data in both places simultaneously.  I didn't initially notice that IIS Logging still existed as a separate feature and worried about the resource impact of such an arrangement.  Fortunately, that is now a non-issue.


~Fin~


Tuesday, October 29, 2013

What's worse than a brick wall?

Trying to build a brick wall without a firm grasp on the concept and practical use of a brick.

I ran into quite the dilemma this evening trying to better understand NetScaler's URL Rewrite/Transform engine.

Really, my dilemma had several facets.
  1. While I have a general idea of how HTTP request and response headers work, this is the first role in which I've been expected to manipulate them.
  2. Learning from the Citrix eDocs site how the specific machinations of a particular product feature works is easily done.  Learning the context in which such machinations are or should be used - not so much.
  3. I didn't actually know what to call - and so, how to search for - the information I needed, so I breadcrumbed all over the goddamned Internet only to end up back on Citrix's eDocs site a few URLs away from where I started.
I slammed headlong into these problems because I wanted to know how to accomplish some URL rewriting wizardry on the NetScaler 7500 appliance(s) we have at work. 

I know now that the NetScaler's system for rewriting URL's is evidently proprietary (though its core web services are built on Apache), and that understanding how to achieve what I want demands I understand (and type into a Google box for) "NetScaler http rewrite expressions" for versions 9.3 and 10.1. 

Obvious in hindsight.  Shut your whore mouth.

What I have:

  • Several IIS 7.5 servers will host the production instance of our much-ado new public web portal.  I will call these servers WEB01, WEB02, and WEB03, and all three servers run the same two web applications:
    • The "main" portal web application, accessed via a client's browser
    • The mobile application, developed and boxed by the vendor to handle connections from iOS and Android applications.
  • Sitting in front of my servers is, as mentioned, a NetScaler 7500.  On this appliance are the server and service objects, as well as the single LB VS object that internally represent my three web servers.
  • There are two domain names that will represent both applications across my three servers to the world: one for external use, one for internal use.  They, with their mobile URL in tow, look something like this:
    • https://site.publicdomain.com
    • https://site.publicdomain.com/mobile
    • https://site.internaldomain.com
    • https://site.internaldomain.com/mobile
  • I have two SSL certificates in play here, one for each namespace.  The certificate for my publicdomain.com namespace comes from a publicly trusted CA, while the internaldomain.com cert comes from the enterprise CA (which I think is just an ADCS server).
Regardless of this monster's final (back-end) form, the basic flow of things will go something like this:
  1. A client types https://site.publicdomain.com into their favorite browser.
  2. Because the public DNS is bound to a publicly accessible VIP on the NetScaler, our beloved NetScaler 7500 receives the request.
  3. Because this is a secure site, a secure connection is negotiated using the public SSL cert that's sitting in the NetScaler's internal certificate store.
  4. The NetScaler makes a load balancing decision based on the status of the three web servers its virtual server is configured to monitor and picks WEB02.
  5. The NetScaler negotiates a secure connection with WEB02 using the enterprise-issued SSL certificate sitting in its internal certificate stores.
  6. Client requests and server responses are moved back and forth, and all's well in my world.
Firstly, I'm going to take a moment to justify #5.  Much of the web traffic I'm administering is sensitive.  Highly so.  Like, medical record kinda' shit.  While there are a dozen good arguments for terminating SSL completely at the NetScaler and connecting unencrypted to the (presumably trusted) back-end, our back-end is not particularly trustworthy.

<pause for gratuitous innuendo>

Secondly, it is in steps 4-6 that one of my major points of indecision comes to haunt me.  If I simply pass the REQ headers to WEB02 as-is, then WEB02 must be configured to answer for any potentially requested host.  That is, it must have site bindings for both the public and internal namespaces, OR it must have a single site binding configured to use a SAN-enabled SSL certificate.

That "OR" part just hit me.  Going into that thought I believed there existed the need for two separate bindings - one for each namespace.  That would require two server certificates, and two IP's, since only one cert can be bound to an IP (without some of Server 2012's new hotness).  I could instead simply have one SSL certificate issued with a Subject and Subject Alternate Name to represent both namespaces and bound to a single site.

The alternative to passing the request's HOST header through is to rewrite it from the public namespace to the internal one, and only have to configure the web servers to answer for the one hostname.  Of course even this alternative belies an apparent obsession with locking down a site to answer only for a specific host.  I'm not 100% sure from whence this obsession stems, but ... there it is.

I should stop now.  Plenty more tomorrow.  This will all make sense by my third cup of coffee and after my second axe murder.

(J/K on the axe murder, Friendly NSA Analyst)

~Fin~

Sunday, October 27, 2013

The idea came in the most surprising of packages

 I don't typically "just run."

I'm not a runner.  I engage in physical activities, sure, but this morning's decision to run around the neighborhood park was atypical.  This past week saw my attendance at my local Crossfit gym (or, as we say in the Legion, the "box") lacking, and I suppose some primal part of me determined it needed the satisfaction brought only by physical discomfort, sweat, and mediocre breathing technique.

Leading up to my first footsteps on the (literally) broken path to the park, my thought process went something like this:
  1. I feel like a fatty.  I should go running.
  2. Why the hell do I have so many more workout shirts than shorts?  Where did all my shorts go?
  3. I should wear a hoodie.  Nothing is worse than cold-air-runner's-lungs.
  4. It's 58F out.  I should not wear a hoodie.
  5. I don't believe the manufacturer of these "no show" athletic socks anticipated how low cut crossfit shoes actually are.
Dressed, stretched, and now walking - making what must appear to be ridiculous windmill motions with my arms as I go - I decide to focus my mind on work.  It is partly a distraction, to keep from dwelling upon the physical discomfort I'm soon to experience, but also a hope that the "experience" of the run will allay fears of the coming week and sharpen any lingering ideas I have about any manner of web-administer-y topics.

Somewhere around 150bpm, only a few minutes into the run, one idea revealed itself above and before all others.

I should blog about this shit.

Why?

I have never been a server administrator before.  Sure, I've studied and certified on topics of server administration, and I have through those studies and personal experiences with other administrators formulated a host of ideas of how such a thing is or isn't, should work or shouldn't.  However, until my current job, it was all armchair quarterbacking.  It was my puerile attempt to sit at the adult's table and talk about something more than just how slow Internet Explorer was, or how someone's "Internet is broken."

Now I'm at the adult's table.


I have metric crap-tons of new material to learn, new ideas to explore, new swords on which to hurl myself unknowingly and otherwise, and a whole host of mistakes to make.  I decided I needed a vessel to receive my thoughts on these things, a medium on which I could present what I learn, what I accidentally set fire and leave alone to crash into the sea.

I am an IIS Administrator.  This is my story.

~Fin~