Sunday, May 18, 2008

WPAR and LPAR comparison

IBM has taken a leadership role in innovation, over the past fourteen years and
has been number one in the patent technology race. Out of this has come a
plethora of new and innovative products. In 2001 IBM announced the LPAR
feature on IBM eServer pSeries and then in 2004 Advanced Power Virtualization
provided the micropartitioning feature. In 2007, IBM announces WPAR Mobility.
WPARs are not a replacement to LPARs. These two technologies are both key
components of IBM's virtualization strategy. The two technologies are
complementary, and can be used together to extend their individual values.
Providing both LPAR and WPAR technology offers a broad range of virtualization
choices to meet the ever changing needs in the IT world. Table 2-2 compares
and contrasts the benefits of the two technologies.
Table 2-2 Comparing WPAR and LPAR
Workload Partitions Logical Partitions
WPARs share OS images LPARs execute OS images
Finer-grained resource management,
per-workload
Resource management per LPAR
Capacity on demand
Security isolation Stronger security isolation
Easily shared files and applications Supports multiple OSes, Tunable to
applications
Lower administrative costs:
1 OS to manage
Easy create/destroy/configure
Integrated management tools
OS Fault isolation
Chapter 2. Understanding and Planning for WPARs 37
Draft Document for Review August 6, 2007 12:52 pm 7431CH_TECHPLANNING.fm
Figure 2-7 shows how LPAR and WPARs can be combined within the same
physical server, which also hosts the WPAR Manager and NFS server required to
support partition mobility.
Important: When considering the information in Table 2-2 you should keep in
mind the following guidelines:
In general, when compared to WPARs, LPARs will provide a greater
amount of flexibility in supporting your system virtualization strategies.
Once you have designed an optimal LPAR resourcing strategy, then within
that strategy you design your WPAR strategy to further optimize your
overall system virtualization strategy in support of AIX6 applications. See
Figure 2-7 for an example of this strategy, where multiple LPARs are
defined to support different OS and application hosting requirements, while
a subset of those LPARs running AIX6 are setup specifically to provide a
global environment for hosting WPARs.
Because LPAR provisioning is hardware/firmware based you should
consider LPARs as a more secure starting point for meeting system
isolation requirements than WPARs.
7431CH_TECHPLANNING.fm Draft Document for Review August 6, 2007 12:52 pm
38 Workload Partitions in IBM AIX Version 6.1

AIX Monitoring


Monitoring AIX Made Easy

Applications Manager monitors the performance of IBM AIX Systems. First, Applications Manager discovers each AIX machine and then monitors the CPU activity, complete memory utilization, and local and remote system statistics.


The AIX Management feature optimizes AIX system performance, delivers comprehensive management reports and ensures availability through automated event detection and correction. Applications Manager also monitors processes running in the AIX system.

Some of the components that are monitored in IBM AIX are:

CPU Utilization Monitor CPU usage - check if CPUs are running at full capacity or are they being underutilized.
Memory Utilization Avoid the problem of your windows system running out of memory. Get notified when the memory usage is high (or memory is dangerously low).
Disk I/O Stats specifies read/writes per second, transfers per second, for each device.
Disk Utilization Maintain a margin of available disk space. Get notified when the disk space falls below the margin. You can also run your own programs/scripts to clear disk clutter when thresholds are crossed.
Process Monitoring Monitor critical processes running in your system. Get notified when a particular process fails.

IBM AIX Monitoring Capabilities
Out-of-the-box management of IBM AIX availability and performance.
Monitors performance statistics such as CPU utilization, memory utilization, disk utilization, Disk I/O Stats and response time.
Mode of monitoring includes Telnet and SSH.
Monitors processes running in AIX systems.
Based on the thresholds configured, notifications and alerts are generated if the AIX system or any specified attribute within the system has problems. Actions are executed automatically based on configurations.
Performance graphs and reports are available instantly. Reports can be grouped and displayed based on availability, health, and connection time.
Delivers both historical and current AIX performance metrics, delivering insight into the performance over a period of time.
Monitors memory usage and detects top consumers of memory.
For more information, refer to IBM AIX Monitoring Online Help.

Database Monitoring

Database Management - Made Easy

Applications Manager is a Database Server monitoring tool that can help monitor a heterogeneous database server environment that may consist of Oracle database, MS SQL, Sybase, IBM DB2 and MySQL databases. It also helps database administrators (DBAs) and system administrators by notifying about potential database performance problems. For database server monitoring, Applications Manager connects to the database and ensures it is up. Applications Manager is also an agentless monitoring tool that executes database queries to collect performance statistics and send alerts, if the database performance crosses a given threshold. With out-of-the box reports, DBAs can plan inventory requirements and troubleshoot incidents quickly.

Database Server Monitoring Software Needs to
ensure high availability of database servers
keep tab on the database size, buffer cache size, database connection time
analyze the number of user connections to the databases at various time
analyze usage trends
help take actions proactively before critical incidents occur.
Applications Manager supports the monitoring of the following databases out-of-the-box:

Oracle Management

MySQL Management

Sybase Management

MS SQL Management

DB2 Management


Oracle Management

Oracle Monitoring includes efficient and complete monitoring of performance, availability, and usage statistics for Oracle databases. It also includes instant notifications of errors and corrective actions. Provides comprehensive reports and graphs. More on Oracle Management >>

MySQL Management

MySQL is the most popular open source relational database system. Applications Manager MySQL Monitoring includes managing MySQL as part of your IT infrastructure, by diagnosing performance problems in real time. More on MySQL Management >>

MS SQL Server Management

Microsoft SQL Server is the enterprise database solution used most commonly on Windows. Applications Manager manages MS SQL Server databases through native Windows performance management interfaces. This ensures optimal and complete access to all the metrics that MS SQL Server exposes. More on MS SQL Management>>

DB2 Management

DB2 Monitoring includes effective monitoring of availability and performance of DB2 Databases with ease. Applications Manager facilitates automated and on-demand monitoring tasks, which will help manage DB2 databases running at its highest levels of performance. More on DB2 Management>>

Sybase Management

Availability and Performance of Sybase ASE Database servers are monitored by Applications Manager. Performance Metrics such as memory usage, connection statistics, etc. are monitored More on Sybase Management>>

Oracle Management

Take Control of Oracle Monitoring



Most business critical applications are database driven. The Oracle database management capability helps database administrators to seamlessly detect, diagnose and resolve Oracle performance issues and monitor Oracle 24X7. The database server monitoring tool is an agentless monitoring software that provides out-of-the-box performance metrics and helps you visualize the health and availability of an Oracle Database server farm. Database administrators can login to the web client and visualize the status and Oracle performance metrics.


Applications Manager also provides out-of-the-box reports that help analyze the database server usage, Oracle database availability and database server health.

Additionally the grouping capability helps group your databases based on the business process supported and helps the operations team to prioritize alerts as they are received.

Some of the components that are monitored in Oracle database are:

Response Time
User Activity
Status
Table Space Usage
Table Space Details
Table Space Status
SGA Performance
SGA Details
SGA Status
Performance of Data Files
Session Details
Session Waits
Buffer Gets
Disk Reads
Rollback Segment


Note: Oracle Application Server performance monitoring is also possible in Applications Manager.
Oracle Management Capabilities
Out-of-the-box management of Oracle availability and performance.
Monitors performance statistics such as user activity, status, table space, SGA performance, session details, etc. Alerts can be configured for these parameters.
Based on the thresholds configured, notifications and alerts are generated. Actions are executed automatically based on configurations.
Performance graphs and reports are available instantly. Reports can be grouped and displayed based on availability, health, and connection time.
Delivers both historical and current Oracle performance metrics, delivering insight into the performance over a period of time.

WebSphere Admin Console

Many a thanks to Satya Dinesh Babu Manne, one of our customers who had found a new way to troubleshoot websphere problem. The solution [What he has basically tried was instead of trying to reuse any existing ports which seem to be having some conflicts, he has defined some new ports and transport chains] is given below:

1) In WebSphere Admin Console, Navigate to Application Servers -> Server Name -> Web Container Settings -> Web Container Transport Chains
2) In this view which shows current transport chains, click on New Button
3) In the resulting wizard at step 1, Give a new name to this chain (I gave it WC_CacheMonitor_Inbound) , and from the template Drop Down box select Webcontainer (Chain 1) and click on Next

4) In Step 2 , give this a new port name to identify it , and the host , port values, For the Port I gave 9030 when creating on instance 1 and 9032 when creating on instance 2. Click on Next.
5) In Step 3, Click on Finish button.
6) Repeat the above steps for each server in Cluster (I got 4 servers)
7) Save Configuration Changes.
Navigate to Environment -> Virtual Hosts, Click on New button
9) In the Wizard, give a new name and click on OK button.
10) In the resulting window click on the new Virtual Host created and click on Host Aliases for that Virtual Host.

11) Add the Virtual Host by making sure to reflect the Host and Port numbers (like 9030, 9032 etc) which have been already been created in the previous steps for Web Container Transport chains.
12) Save the Configuration Changes.
13) Navigate to Applications -> Enterprise Applications -> perfServletApp –> Map virtual hosts for Web modules
14) Select the newly created Virtual Host from the Drop Down.
15) Save the Configuration Changes, and restart all Servers.
16) The perfservlet is now accessible though ports 9030 and 9032 against the hosts configured

I was able to configure and test a websphere monitor after making these changes.

HACMP

High Availability Cluster Multi-Processing (HACMP™) on Linux is the IBM tool for building Linux-based computing platforms that include more than one server and provide high availability of applications and services.

Both HACMP for AIX and HACMP for Linux versions use a common software model and present a common user interface (WebSMIT). This chapter provides an overview of HACMP on Linux and contains the following sections:
•Overview
•Cluster Terminology
•Sample Configuration with a Diagram
•Node and Network Failure Scenarios
•Where You Go from Here.

Overview:

HACMP for Linux enables your business application and its dependent resources to continue running either at its current hosting server (node) or, in case of a failure at the hosting node, at a backup node, thus providing high availability and recovery for the application.
HACMP detects component failures and automatically transfers your application to another node with little or no interruption to the application’s end users.
HACMP for Linux takes advantage of the following software components to reduce application downtime and recovery:
•Linux operating system (RHEL or SUSE ES versions)
•TCP/IP subsystem
•High Availability Cluster Multi-Processing (HACMP™) on Linux cluster management subsystem (the Cluster Manager daemon).
HACMP for Linux Cluster Overview
Overview
14 HACMP for Linux: Installation and Administration Guide
1

HACMP for Linux provides:

•High Availability for system processes, services and applications that are running under HACMP’s control. HACMP ensures continuing service and access to applications during hardware or software outages (or both), planned or unplanned, in an eight-node cluster. Nodes may have access to the data stored on shared disks over an IP-based network (although shared disks cannot be part of the HACMP for Linux cluster and are not kept highly available by HACMP).
•Protection and recovery of applications when components fail. HACMP protects your applications against node and network failures, by providing automatic recovery of applications.
If a node fails, HACMP recovers applications on a surviving node. If a network or a network interface card (adapter) fails, HACMP uses an alternate networks, an additional network interface or an IP label alias to recover the communication links and continue providing access to the data.
•WebSMIT, a web-based user interface to configure an HACMP cluster. In WebSMIT, you can configure a basic cluster with the most widely used, default settings, or configure a customized cluster while having the access to customizable tools and functions. WebSMIT lets you view your existing cluster configuration in different ways (node-centric view, or application-centric view) and provides cluster status tools.
•Easy customization of how applications are managed by HACMP. You can configure HACMP to handle applications in the way you want:
•Applications startup.You select from a set of options for how you want HACMP to start up applications on the node(s).
•Applications recovery actions that HACMP takes. If a failure occurs with an application’s resource that is monitored by HACMP, you select whether you want HACMP to recover applications on another cluster node, or stop the applications.
•HACMP’s follow-up after recovery. You select how you want HACMP to react in cases when you have restored a failed cluster component. For instance, you decide on which node HACMP should restart the application that was previously automatically stopped (or moved to another node) due to a previously detected resource failure.
•Built-in configuration, system maintenance and troubleshooting functions. HACMP has functions to help you with your daily system management tasks, such as cluster administration, automatic cluster monitoring of the application’s health, or notification upon component failures.
•Tools for creating similar clusters from an existing “sample” cluster. You can save your existing HACMP cluster configuration in a cluster snapshot file, and later recreate it in an identical cluster in a few steps.

Cluster Terminology

The list below includes basic terms used in the HACMP environment.
Note:In general, terminology for HACMP is based on industry conventions for high availability. However, the meaning of some of the terms in HACMP may differ from the generic terms.
An application is a service, such as a database, or a collection of system services and their dependent resources, such as a service IP label and application’s start and stop scripts, that you want to keep highly available with the use of HACMP.
An application server is a collection of application start and stop scripts that you provide to HACMP by entering the pathnames for the scripts in the WebSMIT user interface. An application server becomes a resource associated with an application, you include it in a resource group for HACMP to keep it highly available. HACMP ensures that the application can start and stop successfully no matter on which cluster node it is being started.
A cluster node is a physical machine, typically an AIX or a Linux server on which you install HACMP. A cluster node also hosts an application. A cluster node serves as a server for application’s clients. HACMP’s role is to ensure continuous access to the application, no matter on which node in the cluster the application is currently active.
A home node is a node on which the application is hosted, based on your default configuration for the application’s resource group, and under normal conditions.
A takeover node is a backup cluster node to which HACMP may move the application. You can move the application to this node manually, for instance, to free the home node for planned maintenance. Or, HACMP moves the application automatically, due to a cluster component failure.
In HACMP for Linux v.5.4.1, a cluster configuration includes up to eight nodes. Therefore, you can have more than one potential takeover nodes for a particular application. You define the list of nodes on which you want HACMP to host your application using the WebSMIT interface. This list is called a resource group’s nodelist.
A cluster IP network is used for cluster communications between the nodes and for sending heartbeating information. All IP labels configured on the same HACMP network share the netmask, but may be required to have different subnets.
An IP label is a name of a network interface card (NIC) that you provide to HACMP. Network configuration for HACMP requires planning for several types of IP labels:
•Base (or boot) IP labels on each node—the ones through which an initial cluster connectivity is established.
•Service IP labels for each application—the ones through which a connection for a highly available application is established.
•Backup IP labels (optional).
•Persistent IP labels on each node. These are node-bound IP labels that are useful to have in the cluster for administrative purposes.

Note that to ensure high availability and access to the application, HACMP “recovers” the service IP address associated with the application on another node in the cluster in cases of network interface failures. HACMP uses IP aliases for HACMP networks. For information, see Planning IP Networks and Network Interfaces.
An IP alias is an alias placed on an IP label. It coexists on an interface along with the IP label. Networks that support Gratuitous ARP cache updates enable configuration of IP aliases.
IP Address Takeover (IPAT) is a process whereby a service IP label on one node is taken over by a backup node in the cluster. HACMP uses IPAT to provide high availability of IP service labels that belong to resource groups. These labels provide access to applications. HACMP uses IPAT to recover the IP label on the same node or the backup node. HACMP for Linux by default supports the mode of IPAT known as IPAT via IP Aliasing. (The other method of IPAT—IPAT via IP Replacement is not supported).
IP Address Takeover via IP Aliasing is the default method of IPAT used in HACMP. HACMP uses IPAT via IP Aliasing in cases when it must automatically recover a service IP label on another node. To configure IPAT via IP Aliasing, you configure service IP labels and their aliases to the system. When HACMP performs IPAT during automatic cluster events, it places an IP alias recovered from the “failed” node on top of the service IP address on the takeover node. As a result, access to the application continues to be provided.
Cluster resources can include an application server and a service IP label. All or some of these resources can be associated with an application you plan to keep highly available. You include cluster resources into resource groups.
A resource group is a collection of cluster resources.
Resource group startup is an activation of a resource group and its associated resources on a specified cluster node. You choose a resource group startup policy from a predefined list in WebSMIT.
Resource group fallover is an action of a resource group, when HACMP moves it from one node to another. In other words, a resource group and its associated application fall over to another node. You choose a resource group fallover policy from a predefined list in WebSMIT.
Takeover is an automatic action during which HACMP takes over resources from one node and moves them to another node. Takeover occurs when a resource group falls over to another node. A backup node is referred to as a takeover node.
Resource group fallback is an action of a resource group, when HACMP returns it from a takeover node back to the home node. You choose a resource group fallback policy from a predefined list in WebSMIT.
Cluster Startup is the starting of HACMP cluster services on the node(s).
Cluster Shutdown is the stopping of HACMP cluster services on the node(s).
Pre- and post-events are customized scripts provided by you (or other system administrators), which you can make known to HACMP and which will be run before or after a particular cluster event. For more information on pre- and post-event scripts, see the chapter on Planning Cluster Events in the HACMP for AIX Planning Guide.