The CCNA certification is valid for 3 years and mine was due to expire at the end of August 2012. I could either retake the same exam and recertify, or take another CCNA "concentration" exam that would give me a new certification and renew the original certification at the same time. I opted to tackle the CCNA Security exam, IINS 640-553.
I'd originally bought the Cisco Press "Authorized Self-Study Guide", Implementing Cisco IOS Network Security by Catherine Paquet back in 2010, but the material is a bit dry and I didn't have the motivation to get into it very far. By booking the exam, I suddenly acquired the motivation required.
As things happen, the 640-553 exam is due to be retired in September 2012, to be replaced by 640-554. The main difference in the new exam appears to be an additional focus on the Cisco ASA platform, as well as de-emphasising the Cisco Secure Device Manager (SDM). This means that any advice I give here will be redundant soon, and also I'm bound by the NDA, so can't obviously comment on what is in the exam.
What I can do though is give some general thoughts on the revision process:
The Good
The Implementing Cisco IOS Network Security book is very thorough. It covers a lot of detail and assumes little prior knowledge of security. Some of it is dry, especially the first chapter which weighs in at about 100 pages and gives an introduction to the world of security. Once that's passed, the content gets better and even the chapter on cryptography was interesting(!).
I also bought the Cisco CCNA Security Lab manual for the CCNA Security course. This gave some very good exercises to run through which were very useful in grounding the theory in the practical.
All of this was made possible using the amazing GNS3 router simulation software. I installed this on a meaty Windows Server VM and was able to run the 3 routers and 2 XP images (in Virtualbox, under ESXi) without any problems. The ability to save configurations and easily re-import them was a great time saver. GNS3 doesn't do everything (specifically switches, due to the custom silicon in them), but it made the whole process of learning the syllabus a lot easier.
There is some very good material at the Cisco Learning Network including free study chapters, training videos and discussions. Highly recommended.
The Bad
Cisco sell the book for self-study, but make it very difficult to practice because IOS images are not available without having the correct support contract. If you work for a large company with either old routers sat on a shelf or a contract with the ability to download the image then you'll be okay. Otherwise I guess you'll be searching the Internet for a dodgy copy of an old image. Seriously Cisco, how about making them freely available? You can do the study material with a 2600 series router and how old is that?
The same is true of the IPS signatures. A valid contract is required just to learn how the IPS works and again, this could mean a trip to the darker parts of the Internet to find them.
The Cisco Press book covers the Cisco Access Control Server software but it's not in the syllabus or lab manual. It can be used to learn about AAA and specifically authentication and authorization with RADIUS and TACACS+. Unfortunately Cisco don't have a trial version to help self-studying students.
The Ugly
Getting the Cisco Security Device Manager (SDM) working requires jumping through a number of hoops. To cut a long story short, you need an old version of Java (1.4 worked for me) and Windows XP. I'm guessing the latter requirement is due to Internet Explorer 6 as I couldn't get it working on Server 2008 R2 no matter what settings I tried.
Conclusion
Having worked through the labs a number of times and then setting things up "blind" (without referring to any notes), I felt fairly confident as I went into the exam. I passed with a good mark well above the passing level, so I'm naturally very pleased with this. It's a good subject to read up on since security requirements impact on so much of what we design these days. The CCNA Security should demonstrate I now have a solid grounding in the subject, even if I'm still a long way from being an expert.
Wednesday, 8 August 2012
Tuesday, 10 July 2012
Home lab upgrade: HP ML110 G7
My home vSphere infrastructure has, until very recently, consisted of:
Unfortunately, the ML115 G5 PSU died recently, leaving me with a single server. While looking at the options (the PSU was really expensive to replace!), I came across the ML110 G7 cash back offer (£150!). Having counted the pennies and done some reading up, I ordered two of these to replace the existing ML servers. I replaced the included 2GB DIMM with 4x4GB DIMMs. (side question: what do people do with the 1GB and 2GB DIMMs that come with servers that we immediately replace with something bigger???).
This arrived today:
The inside of the ML110 G7 is pretty clean and fitting the extra RAM was very straightforward:
The G7 is roughly the same size as the G5:
And with both of them in place (with the Cisco SG200-26 and Belkin SOHO 4 port KVM on top):
In order to make the ML110 G7 work with VMware ESXi, the BIOS needs to be configured with the correct "C State". I used this very useful blog post to set mine.
I installed ESXi using a USB CD-ROM drive and everything went very smoothly. Added both to vCenter Server, mounted an NFS datastore, setup iSCSI and configured the vSwitch for correct VLAN membership and everything is now ready to go. This upgrade means I'll be in a better position to test vCloud Director, vShield etc.
- 2 x HP Microserver N36L
- 1 x HP ML110 G5
- 1 x HP ML115 G5
Unfortunately, the ML115 G5 PSU died recently, leaving me with a single server. While looking at the options (the PSU was really expensive to replace!), I came across the ML110 G7 cash back offer (£150!). Having counted the pennies and done some reading up, I ordered two of these to replace the existing ML servers. I replaced the included 2GB DIMM with 4x4GB DIMMs. (side question: what do people do with the 1GB and 2GB DIMMs that come with servers that we immediately replace with something bigger???).
This arrived today:
The inside of the ML110 G7 is pretty clean and fitting the extra RAM was very straightforward:
The G7 is roughly the same size as the G5:
And with both of them in place (with the Cisco SG200-26 and Belkin SOHO 4 port KVM on top):
In order to make the ML110 G7 work with VMware ESXi, the BIOS needs to be configured with the correct "C State". I used this very useful blog post to set mine.
I installed ESXi using a USB CD-ROM drive and everything went very smoothly. Added both to vCenter Server, mounted an NFS datastore, setup iSCSI and configured the vSwitch for correct VLAN membership and everything is now ready to go. This upgrade means I'll be in a better position to test vCloud Director, vShield etc.
Tuesday, 29 May 2012
Mac Mini upgrade
My main computer is a mid-2007 Mac Mini. While it's performed very well in the time I've had it, recently I've become frustrated with how slow it is running with Mac OS X Lion.
The cost of a new Mac is beyond what I have to spend at the moment, so I looked into ways to upgrade the Mini. I've previously been hesitant to do this as getting inside the Mini requires some effort. But these are difficult times, so I bit the bullet and opted to upgrade the memory and replace the internal hard disk.
RAM upgrade
I wanted to upgrade the RAM from the 2GB that it came with. This is the maximum that Apple officially support, but you can put more in and it will work. There is a caveat: Due to the chipset used, only 3GB can be addressed by the operating system (even though it will detect more). This meant I could either go for 1GB + 2GB solution, or 2GB + 2GB and effectively waste 1GB of it. I eventually went for the latter because I had read that putting equal sized DIMMs enabled the RAM to run at full speed. I don't know if this is really the case, but the cost difference was very small.
The RAM I bought from Ebuyer was the Crucial 2GB DDR2 667Mhz/PC2-5300 Laptop Memory SODIMM CL5 1.8V (quickfind code: 142421) and cost just over £20 per stick. I bought two sticks.
Disk upgrade
The hard disk that came in the Mac Mini was a 120GB SATA 5400RPM drive. The rotational speed of the disk is at the low end (typical desktops are 7200RPM, with servers having 10000RPM or 15000RPM drives). It was never particularly speedy, but did the job. With the cost of SSDs falling, I could replace the disk with something solid state and much faster.
Ebuyer had an offer on the OCZ 120GB Agility 3 SSD (AGT3-25SAT3-120G) (quickfind code: 268244) costing £80.
There is a lot of discussion about the support for TRIM in Mac OS X. TRIM is a command that the operating system can send to the device to notify it that disk blocks are free and can be wiped. Without TRIM support, disk writes can slow down over time. Official Apple SSDs support TRIM, third party SSDs don't. There is a kernel extension that hacks third party devices, but that was a bit too risky for my liking.
Reading up on the OCZ Agility 3, it includes on-chipset garbage collection that performs reclamation of blocks in the background. It may not be as good as TRIM, but it does a similar job.
So, I'm running without TRIM and will see over time how it performs.
Performing the upgrade
Due to the way that the Mac Mini is constructed, getting into the case can be a challenge. You will need a putty knife to get inside the case. Once opened, a very small cross-head screwdriver is required to detach the optical and hard drives. Inserting the RAM is straightforward once this is done. Similarly, replacing the hard disk was pretty easy once inside. There are plenty of guides online that show how to do this.
With the internal drive replaced, there is no operating system to boot from. I had burned a DVD image of the Lion installer and used this to boot the machine (hold down "c" when powering up and until the Apple logo and spinning wheel appears). The DVD has the option to restore a Time Machine backup which I selected. You need to load the Disk Utility to format the SSD, at which point Time Machine can be used to restore a backup (held on an external USB drive) to the internal disk.
The restore process started and I walked away. When I came back, the Mini was asleep and I couldn't get it to wake without performing a power off/on. At which point I didn't know whether it had completed or not. I repeated the process and watched it, moving the mouse every couple of minutes in case power management was sending it to sleep. With six minutes of restore remaining, the machine put itself to sleep. This wasn't power management. This was a crash.
I then tried booting off the original Leopard DVD that came with the machine. I performed the Time Machine restore from here and it completed successfully without incident. I'm not sure if this is a bug in the Lion DVD, but it's worth noting if you have the same problem!
With the machine restored, I booted off the SSD and loaded a few apps. The extra memory and SSD make it feel like a new computer! Where the original configuration was taking up to 2 minutes from power on to the desktop, it now does it in 48 seconds (including Finder opening network drives)! Most applications load very quickly with no noticeable delay.
I previously used Google Mail through a browser as Mail.app felt like unnecessary overhead and was quite slow. Now it's fast and I can keep it open all the time in the background (along with iCal which was another application I had given up on).
Spending £120 on a couple of upgrades has given the Mac Mini a new lease of life. Although the hardware won't be up to running Mountain Lion when it's out later this year, I should be able to get another couple of years out of the Mini, which is definitely worth it.
The cost of a new Mac is beyond what I have to spend at the moment, so I looked into ways to upgrade the Mini. I've previously been hesitant to do this as getting inside the Mini requires some effort. But these are difficult times, so I bit the bullet and opted to upgrade the memory and replace the internal hard disk.
RAM upgrade
I wanted to upgrade the RAM from the 2GB that it came with. This is the maximum that Apple officially support, but you can put more in and it will work. There is a caveat: Due to the chipset used, only 3GB can be addressed by the operating system (even though it will detect more). This meant I could either go for 1GB + 2GB solution, or 2GB + 2GB and effectively waste 1GB of it. I eventually went for the latter because I had read that putting equal sized DIMMs enabled the RAM to run at full speed. I don't know if this is really the case, but the cost difference was very small.
The RAM I bought from Ebuyer was the Crucial 2GB DDR2 667Mhz/PC2-5300 Laptop Memory SODIMM CL5 1.8V (quickfind code: 142421) and cost just over £20 per stick. I bought two sticks.
Disk upgrade
The hard disk that came in the Mac Mini was a 120GB SATA 5400RPM drive. The rotational speed of the disk is at the low end (typical desktops are 7200RPM, with servers having 10000RPM or 15000RPM drives). It was never particularly speedy, but did the job. With the cost of SSDs falling, I could replace the disk with something solid state and much faster.
Ebuyer had an offer on the OCZ 120GB Agility 3 SSD (AGT3-25SAT3-120G) (quickfind code: 268244) costing £80.
There is a lot of discussion about the support for TRIM in Mac OS X. TRIM is a command that the operating system can send to the device to notify it that disk blocks are free and can be wiped. Without TRIM support, disk writes can slow down over time. Official Apple SSDs support TRIM, third party SSDs don't. There is a kernel extension that hacks third party devices, but that was a bit too risky for my liking.
Reading up on the OCZ Agility 3, it includes on-chipset garbage collection that performs reclamation of blocks in the background. It may not be as good as TRIM, but it does a similar job.
So, I'm running without TRIM and will see over time how it performs.
Performing the upgrade
Due to the way that the Mac Mini is constructed, getting into the case can be a challenge. You will need a putty knife to get inside the case. Once opened, a very small cross-head screwdriver is required to detach the optical and hard drives. Inserting the RAM is straightforward once this is done. Similarly, replacing the hard disk was pretty easy once inside. There are plenty of guides online that show how to do this.
With the internal drive replaced, there is no operating system to boot from. I had burned a DVD image of the Lion installer and used this to boot the machine (hold down "c" when powering up and until the Apple logo and spinning wheel appears). The DVD has the option to restore a Time Machine backup which I selected. You need to load the Disk Utility to format the SSD, at which point Time Machine can be used to restore a backup (held on an external USB drive) to the internal disk.
The restore process started and I walked away. When I came back, the Mini was asleep and I couldn't get it to wake without performing a power off/on. At which point I didn't know whether it had completed or not. I repeated the process and watched it, moving the mouse every couple of minutes in case power management was sending it to sleep. With six minutes of restore remaining, the machine put itself to sleep. This wasn't power management. This was a crash.
I then tried booting off the original Leopard DVD that came with the machine. I performed the Time Machine restore from here and it completed successfully without incident. I'm not sure if this is a bug in the Lion DVD, but it's worth noting if you have the same problem!
With the machine restored, I booted off the SSD and loaded a few apps. The extra memory and SSD make it feel like a new computer! Where the original configuration was taking up to 2 minutes from power on to the desktop, it now does it in 48 seconds (including Finder opening network drives)! Most applications load very quickly with no noticeable delay.
I previously used Google Mail through a browser as Mail.app felt like unnecessary overhead and was quite slow. Now it's fast and I can keep it open all the time in the background (along with iCal which was another application I had given up on).
Spending £120 on a couple of upgrades has given the Mac Mini a new lease of life. Although the hardware won't be up to running Mountain Lion when it's out later this year, I should be able to get another couple of years out of the Mini, which is definitely worth it.
Friday, 18 May 2012
Personal Cloud: What I use (May 2012 edition)
This blog originally started because I was looking for ways to move much of my online life to cloud services. It's since grown to encompass some of my technical projects, but I still have an interest in what is now referred to as "personal cloud" services. So, as of May 2012, these are the services I use:
Email
I use Google Mail for my main, personal email. For signing up to web sites etc., I also have a Yahoo email address which supports disposable addresses, very useful for creating per-domain, unique addresses that can be removed if I start getting spam. Just for completeness, I also have a Hotmail account which gets limited use and is used primarily for logging into Microsoft services.
Calendar
Google Calendar synchronises with my iPhone/iPad and my Outlook calendar using Google Calendar Sync.
News
Google Reader remains my main application for RSS feed aggregation. I use it with Reeder on the iPhone and iPad for mobile reading.
Notes
I absolutely love Evernote. I use it to store multiple notebooks containing all my research, notes, web clipping (especially useful now Evernote Clearly has been released) and it acts as a single "dumping ground" in which to throw my thoughts and anything I find interesting. I run the client on Windows, Mac, iPhone and iPad. It's so good I pay for it as a premium user.
Backups
I have a Crashplan+ subscription and use it to backup my Windows PC, T's Windows PC and my Mac. Data is encrypted before leaving my home network so is secure online. There is a real peace of mind knowing that a copy of all our family documents - and photos - are backed up online. It's worth paying for.
Passwords
I've recently started using LastPass to manage all my website passwords as well as provide a place to store other private data (such as computer account credentials). LastPass encrypts data locally, but stores the results in the cloud. There is an app for the iPhone, but you need to be a premium user to get it. I'm only using the free version at the moment, but I like what I've seen with it so far and may subscribe.
[Update: I liked LastPass enough to pay for the Premium version. Recommended!]
File sharing
I've had a Box account for a few years (with 5GB of free space) and it's recently changed its focus to become more of a SharePoint alternative. It's pretty good, but this space is getting crowded with 25GB free with Microsoft SkyDrive, 5GB with Google Drive and 2GB with Dropbox. I wouldn't use any of these services to store my important, personal data (especially since it's unencrypted), but for non-sensitive data, it's good to have options, especially if you need to collaborate with someone. It's too early to determine what service I'll end up focusing on, so watch this space as the products mature.
I use Google Mail for my main, personal email. For signing up to web sites etc., I also have a Yahoo email address which supports disposable addresses, very useful for creating per-domain, unique addresses that can be removed if I start getting spam. Just for completeness, I also have a Hotmail account which gets limited use and is used primarily for logging into Microsoft services.
Calendar
Google Calendar synchronises with my iPhone/iPad and my Outlook calendar using Google Calendar Sync.
News
Google Reader remains my main application for RSS feed aggregation. I use it with Reeder on the iPhone and iPad for mobile reading.
Notes
I absolutely love Evernote. I use it to store multiple notebooks containing all my research, notes, web clipping (especially useful now Evernote Clearly has been released) and it acts as a single "dumping ground" in which to throw my thoughts and anything I find interesting. I run the client on Windows, Mac, iPhone and iPad. It's so good I pay for it as a premium user.
Backups
I have a Crashplan+ subscription and use it to backup my Windows PC, T's Windows PC and my Mac. Data is encrypted before leaving my home network so is secure online. There is a real peace of mind knowing that a copy of all our family documents - and photos - are backed up online. It's worth paying for.
Passwords
I've recently started using LastPass to manage all my website passwords as well as provide a place to store other private data (such as computer account credentials). LastPass encrypts data locally, but stores the results in the cloud. There is an app for the iPhone, but you need to be a premium user to get it. I'm only using the free version at the moment, but I like what I've seen with it so far and may subscribe.
[Update: I liked LastPass enough to pay for the Premium version. Recommended!]
File sharing
I've had a Box account for a few years (with 5GB of free space) and it's recently changed its focus to become more of a SharePoint alternative. It's pretty good, but this space is getting crowded with 25GB free with Microsoft SkyDrive, 5GB with Google Drive and 2GB with Dropbox. I wouldn't use any of these services to store my important, personal data (especially since it's unencrypted), but for non-sensitive data, it's good to have options, especially if you need to collaborate with someone. It's too early to determine what service I'll end up focusing on, so watch this space as the products mature.
Saturday, 21 April 2012
EMC VNXe - diving under the hood (Part 5: Networking)
A quick look at the EMC Community Forum for the VNXe will show a lot of questions around the best way to configure networking. This is partly due to the way that networking is handled by the VNXe.
Put simply, it's different from the CLARiiON and Celerra.
As the previous post illustrated, the VNXe operating system is based on Linux with CSX Containers hosting FLARE and DART environments. To understand how networking is handled in the VNXe requires looking at various parts of the stack.
The physical perspective
Each SP has, by default, four network interfaces:
The Linux perspective
Running ifconfig in an SSH session reveals a number of different devices:
The Linux "mgmt" device maps to the physical management NIC port, but does not have an IP address assigned to it. A virtual interface, "mgmt:0" is created on top of "mgmt" and this is assigned the Unisphere management interface IP address. This is almost certainly due to the HA capabilities built into the VNXe. In the event of a SP failure, the virtual interface will be failed over to the peer SP.
End user data is transferred over "eth2" and "eth3". The first thing to note from the output of ifconfig is that the MAC addresses of both interfaces are the same.
Another device, "bond0", is created on top of eth2. If link aggregation is configured, eth3 is also joined to bond0. This provides load balancing of network traffic into the VNXe.
There is also a "cmin0" device which connects to the internal "CLARiiON Messaging Interface" (CMI). The CMI is a fast PCIe connection to the peer SP and is used for failover traffic and cache mirroring. The cmin0 device does not have an IP address. It's possible the CMI communicates using layer 2 only and therefore doesn't require an IP address, but that's only speculation.
Finally, there is an eth_int device that has an IP address in the 128.221.255.0 subnet. This is used to communicate with the peer SP and either uses the CMI or has an internal network connection of some kind.
A quick check of a client machines ARP cache reveals that the IP addresses of Shared Folder Servers in the VNXe do not have a MAC address that is mapped to any of the Linux devices. So how does IP traffic reach the Shared Folder Servers if the network ports are not listening for those MAC addresses? Running the "dmesg" command shows the kernel log, including boot information. The answer to the question can be found here:
[ 659.638845] device eth_int entered promiscuous mode
[ 659.797273] device bond0 entered promiscuous mode
[ 659.797283] device eth2 entered promiscuous mode
[ 659.799994] device eth3 entered promiscuous mode
[ 659.805381] device cmin0 entered promiscuous mode
On start up, all the Linux network ports are put into promiscuous mode. This means that the ports listen to all traffic passed to them regardless of the destination MAC address and can therefore pass traffic up the stack to the DART container.
The DART perspective
The CSX DART Container sits on top of the Linux operating system and provides its own network devices:
DART also creates some "Fail Safe Network" devices on top of the vnics:
The DART fsnX devices are virtual devices that map to the underlying DART devices in an active/standby configuration:
I don't know what the rep30 interface is, but guess it could be an address for replication to use if licensed and configured.
This might be more clearly explained with a diagram:
Failover
So how does failover work?
It would appear that there are different failover technologies used. The Linux-HA software is used within Linux to provide management interface failover to the peer SP.
It's also likely that DART is doing some form of HA clustering as well. On a Celerra, DART redundancy was handled by setting up a (physical) standby Data Mover. Given that DART is running as a CSX Container, does the peer SP actually run two instances of the DART CSX Container, one active for the SP, the second running standby for the peer? I don't know, but it would make some sense if it did, and would also help explain what the 8GB of RAM in each SP is being used for.
Network Configuration
This document has a good overview on how to best configure networking for a VNXe. The following hopefully explains "why" networks should be configured in a particular way.
The "best" approach does depend on whether you are using NFS or iSCSI. The important thing to understand is that they make use of multiple links in different ways:
Stacking switch pair with link aggregation
If both eth2 and eth3 are aggregated into an Etherchannel, then the failure of one link should not cause a problem. Redundancy is handled at the network layer (through the bond0 device) and DART should not even notice that the physical link is down.
With an aggregation, traffic is load balanced based on a MAC or IP hash. With multiple hosts accessing the VNXe, the load should be balanced pretty evenly across both links. However, if you only have a single host accessing the VNXe, chances are you will be limited to the throughput of a single link. Despite this limitation, you will still have the additional redundancy of the second link.
Separate switches (no link aggregation)
If eth2 and eth3 are connected to separate non-stacking switches then eth3 would not be joined to bond0, but would connect directly to vnic1 and be used by a different Shared Folder Server or iSCSI Server to one using eth2.
Therefore, the only connection on the SP is via eth2 and if it fails, DART would detect that vnic0 has failed and the fsn0 device will failover from vnic0 to vnic0-b. This will then route traffic via the peer SPs physical Ethernet ports via the cmin0 device. Presumably a gratuitous ARP request is sent from the peer SP to notify the upstream network of the new route.
This is why eth0 on SPA must be on the same subnet (and VLAN) as SPB. If a failover occurs, then the peer SP must be able to impersonate the failed link.
This is a big difference from the CLARiiON which passes LUN ownership to the peer SP, or Celerra which relies on standby data movers to pick up the load if the active fails (although as noted above, it might still do this in software). In contrast, there should be little performance hit if traffic is directed across the CMI to the peer SP (although the peer SP network links may be overloaded as a result).
Conclusion
This pretty much concludes this mini-series into the VNXe!
It goes to show that even the simplest of storage devices have a fair amount of complexity under the hood and despite its limitations in some areas, the VNXe is a very good entry level array. It is impressive how EMC have managed to virtualise the CLARiiON and Celerra stacks and it makes sense that this approach will be used in other products in the future.
Thanks for reading! Any comments and/or corrections welcome.
Put simply, it's different from the CLARiiON and Celerra.
As the previous post illustrated, the VNXe operating system is based on Linux with CSX Containers hosting FLARE and DART environments. To understand how networking is handled in the VNXe requires looking at various parts of the stack.
The physical perspective
Each SP has, by default, four network interfaces:
- 1 x Network Management
- 1 x Service Laptop (not discussed here)
- 2 x LAN ports (for iSCSI/CIFS/NFS traffic)
The Linux perspective
Running ifconfig in an SSH session reveals a number of different devices:
- bond0
- cmin0
- eth2
- eth3
- eth_int
- lab (???)
- lo (loopback)
- mgmt
- mgmt:0
The Linux "mgmt" device maps to the physical management NIC port, but does not have an IP address assigned to it. A virtual interface, "mgmt:0" is created on top of "mgmt" and this is assigned the Unisphere management interface IP address. This is almost certainly due to the HA capabilities built into the VNXe. In the event of a SP failure, the virtual interface will be failed over to the peer SP.
End user data is transferred over "eth2" and "eth3". The first thing to note from the output of ifconfig is that the MAC addresses of both interfaces are the same.
Another device, "bond0", is created on top of eth2. If link aggregation is configured, eth3 is also joined to bond0. This provides load balancing of network traffic into the VNXe.
There is also a "cmin0" device which connects to the internal "CLARiiON Messaging Interface" (CMI). The CMI is a fast PCIe connection to the peer SP and is used for failover traffic and cache mirroring. The cmin0 device does not have an IP address. It's possible the CMI communicates using layer 2 only and therefore doesn't require an IP address, but that's only speculation.
Finally, there is an eth_int device that has an IP address in the 128.221.255.0 subnet. This is used to communicate with the peer SP and either uses the CMI or has an internal network connection of some kind.
A quick check of a client machines ARP cache reveals that the IP addresses of Shared Folder Servers in the VNXe do not have a MAC address that is mapped to any of the Linux devices. So how does IP traffic reach the Shared Folder Servers if the network ports are not listening for those MAC addresses? Running the "dmesg" command shows the kernel log, including boot information. The answer to the question can be found here:
[ 659.638845] device eth_int entered promiscuous mode
[ 659.797273] device bond0 entered promiscuous mode
[ 659.797283] device eth2 entered promiscuous mode
[ 659.799994] device eth3 entered promiscuous mode
[ 659.805381] device cmin0 entered promiscuous mode
On start up, all the Linux network ports are put into promiscuous mode. This means that the ports listen to all traffic passed to them regardless of the destination MAC address and can therefore pass traffic up the stack to the DART container.
The DART perspective
The CSX DART Container sits on top of the Linux operating system and provides its own network devices:
- DART vnic0 maps to Linux bond0
- DART vnic1 maps to Linux eth3 (presumably unless eth3 is joined to bond0)
- DART vnic0-b maps to the Linux cmin0 device
- DART vnic-int maps to the Linux eth_int device
DART also creates some "Fail Safe Network" devices on top of the vnics:
- fsn0 maps to vnic0
- fsn1 maps to vnic1
The DART fsnX devices are virtual devices that map to the underlying DART devices in an active/standby configuration:
- fsn1 active=vnic1 primary=vnic1 standby=vnic1-b
- fsn0 active=vnic0 primary=vnic0 standby=vnic0-b
- rep30 - IP on the internal 128.221.255.0 network and maps to vnic-int
- el30 - IP address of DART instance "server_2" on the vnic-int interface, 128.221.255.0 network
- if_12 - maps to device fsn0 and contains the user configured IP address of the Shared Folder Server
I don't know what the rep30 interface is, but guess it could be an address for replication to use if licensed and configured.
This might be more clearly explained with a diagram:
Failover
So how does failover work?
It would appear that there are different failover technologies used. The Linux-HA software is used within Linux to provide management interface failover to the peer SP.
It's also likely that DART is doing some form of HA clustering as well. On a Celerra, DART redundancy was handled by setting up a (physical) standby Data Mover. Given that DART is running as a CSX Container, does the peer SP actually run two instances of the DART CSX Container, one active for the SP, the second running standby for the peer? I don't know, but it would make some sense if it did, and would also help explain what the 8GB of RAM in each SP is being used for.
Network Configuration
This document has a good overview on how to best configure networking for a VNXe. The following hopefully explains "why" networks should be configured in a particular way.
The "best" approach does depend on whether you are using NFS or iSCSI. The important thing to understand is that they make use of multiple links in different ways:
Stacking switch pair with link aggregation
If both eth2 and eth3 are aggregated into an Etherchannel, then the failure of one link should not cause a problem. Redundancy is handled at the network layer (through the bond0 device) and DART should not even notice that the physical link is down.
With an aggregation, traffic is load balanced based on a MAC or IP hash. With multiple hosts accessing the VNXe, the load should be balanced pretty evenly across both links. However, if you only have a single host accessing the VNXe, chances are you will be limited to the throughput of a single link. Despite this limitation, you will still have the additional redundancy of the second link.
Separate switches (no link aggregation)
If eth2 and eth3 are connected to separate non-stacking switches then eth3 would not be joined to bond0, but would connect directly to vnic1 and be used by a different Shared Folder Server or iSCSI Server to one using eth2.
Therefore, the only connection on the SP is via eth2 and if it fails, DART would detect that vnic0 has failed and the fsn0 device will failover from vnic0 to vnic0-b. This will then route traffic via the peer SPs physical Ethernet ports via the cmin0 device. Presumably a gratuitous ARP request is sent from the peer SP to notify the upstream network of the new route.
This is why eth0 on SPA must be on the same subnet (and VLAN) as SPB. If a failover occurs, then the peer SP must be able to impersonate the failed link.
This is a big difference from the CLARiiON which passes LUN ownership to the peer SP, or Celerra which relies on standby data movers to pick up the load if the active fails (although as noted above, it might still do this in software). In contrast, there should be little performance hit if traffic is directed across the CMI to the peer SP (although the peer SP network links may be overloaded as a result).
Conclusion
This pretty much concludes this mini-series into the VNXe!
It goes to show that even the simplest of storage devices have a fair amount of complexity under the hood and despite its limitations in some areas, the VNXe is a very good entry level array. It is impressive how EMC have managed to virtualise the CLARiiON and Celerra stacks and it makes sense that this approach will be used in other products in the future.
Thanks for reading! Any comments and/or corrections welcome.
Thursday, 5 April 2012
EMC VNXe - diving under the hood (Part 4: CSX)
After the last post, I was pointed in the direction of the "VNXe Theory of Operations" online training available from education.emc.com (just do a search for it). This free course provides some interesting details into the VNXe architecture.
With knowledge gained from the course in mind, let's see if we can get a better understanding of what's happening under the hood...
C4LX and CSX
When the VNXe was announced, Chad Sakac at EMC referred to the it as "using a completely homegrown EMC innovation called C4LX and CSX to virtualize, encapsulate whole kernels and other multiple high performance storage services into a tight, integrated package."
In the same blog post, Chad also illustrated the operating system stack which showed the C4LX and CSX components are built on a 64bit Linux kernel.
CSX (short for "Common Software eXecution") is designed to provide a common API layer for EMC software that is not tied to the underlying operating system kernel. As a portable execution environment, CSX can run on many platforms in either kernel or user space. So when some functionality is written within the CSX framework (e.g., data compression), it can be easily ported to all CSX supporting platforms, regardless of whether the underlying operating system is DART, FLARE or something else. Steve Todd has some more details about CSX on his blog.
So if you read that CSX instances are similar to Virtual Machines, think of it in terms more like a Java Virtual Machine rather than a VMware Virtual Machine. It's an API abstraction and runtime environment, not a virtualisation of physical resources such as CPU and memory.
There aren't many details on what C4LX is, but here's my conjecture: There are some functions that CSX needs the underlying operating system to perform that may not be easily possible "out of the box". If that's true, then C4LX is the Linux kernel along with a bunch of kernel modules and additional software that provide this functionality. Or another way to describe it might be to call it EMC's own internal Linux distribution...
Data Path
Like the data plane and control plane in a network switch, software in the VNXe appears to be designed to operate on the "Data Path" or the "Control Path".
CSX creates various "Containers" that are populated with "CSX Modules". A Container is either a user space application or a kernel module. CSX Containers implement functionality within the Data Path.
The FLARE functionality described in part 2 of this series is implemented as a CSX Module, as is the DART functionality described in part 3. Both these modules run in the Linux user space. There is a degree of isolation between Containers in that they be terminated and restarted without interfering with other Containers. However, some Containers (such as DART) have dependencies on other Containers (FLARE).
In addition to the FLARE and DART Containers, a Global Memory Services (GMS) Container provides memory management functionality and services other Containers. As an example, the FLARE Container takes 500MB memory, while the DART Container takes 2.5GB, all allocated by the GMS.
A kernel space Container is responsible for allocating resources on behalf of user space Containers. The Linux Upstart software provides a means to start, stop and restart Containers.
Control Path
The Control Path is a implemented using technology derived from the Celerra Control Station (itself a Linux-based server) and the CLARiiON NaviSphere software. The Control Path is also where the Common Security Toolkit (CST) is found. The CST appears to be RSA technology and is used in multiple EMC products for security-relation functions. In contrast to the Data Path which consists of functionality directly relating to the transferring of data, the Control Path is concerned with management functionality.
Within the Control Path of the VNXe is the EMC CIM (Common Interface Module) Object Manager (ECOM) management server. ECOM interfaces with "Providers" which are essentially plug-ins. Within the VNXe, ECOM runs on the master SP.
There are a number of different Providers. These include Application Providers for Exchange, iSCSI, Shared Folders and VMware software provision, a Virtual Server Provider, Pools Provider, CLARiiON Provider and Celerra Provider. There are also providers for Registration, Scripting, Scheduling, Replication etc. As plug-ins to the ECOM server, additional services can be written to extend the functionality within the VNXe.
With the use of Providers, ECOM implements a middleware subsystem that can be called by front end applications such as Unisphere or the VNXe command line.
In addition to running Providers, ECOM also provides basic web server functionality used by the Unisphere GUI and CLI via the Apache web server.
Pulling it together
The VNXe uses some additional Linux software along with the custom CSX and ECOM components. High availability is implemented through the open source Pacemaker cluster resource manager and using the Softdog software timing kernel driver. CSX components are resource managed using the cgroups feature of the Linux kernel. The Logging system uses the Postgres database. Although this is covered in the EMC training, it's also possible to see this by checking the output of "ps" from an SSH session.
To understand how the various components hang together, the boot sequence looks a bit like this:
Hopefully this gives some insight into the complexity that underpins the VNXe. We're going to look at one more topic to conclude this mini series, and it's a subject that is the source of many questions on the EMC VNXe Community forum. In the next post we'll have a look at VNXe networking...
With knowledge gained from the course in mind, let's see if we can get a better understanding of what's happening under the hood...
C4LX and CSX
When the VNXe was announced, Chad Sakac at EMC referred to the it as "using a completely homegrown EMC innovation called C4LX and CSX to virtualize, encapsulate whole kernels and other multiple high performance storage services into a tight, integrated package."
In the same blog post, Chad also illustrated the operating system stack which showed the C4LX and CSX components are built on a 64bit Linux kernel.
CSX (short for "Common Software eXecution") is designed to provide a common API layer for EMC software that is not tied to the underlying operating system kernel. As a portable execution environment, CSX can run on many platforms in either kernel or user space. So when some functionality is written within the CSX framework (e.g., data compression), it can be easily ported to all CSX supporting platforms, regardless of whether the underlying operating system is DART, FLARE or something else. Steve Todd has some more details about CSX on his blog.
So if you read that CSX instances are similar to Virtual Machines, think of it in terms more like a Java Virtual Machine rather than a VMware Virtual Machine. It's an API abstraction and runtime environment, not a virtualisation of physical resources such as CPU and memory.
There aren't many details on what C4LX is, but here's my conjecture: There are some functions that CSX needs the underlying operating system to perform that may not be easily possible "out of the box". If that's true, then C4LX is the Linux kernel along with a bunch of kernel modules and additional software that provide this functionality. Or another way to describe it might be to call it EMC's own internal Linux distribution...
Data Path
Like the data plane and control plane in a network switch, software in the VNXe appears to be designed to operate on the "Data Path" or the "Control Path".
CSX creates various "Containers" that are populated with "CSX Modules". A Container is either a user space application or a kernel module. CSX Containers implement functionality within the Data Path.
The FLARE functionality described in part 2 of this series is implemented as a CSX Module, as is the DART functionality described in part 3. Both these modules run in the Linux user space. There is a degree of isolation between Containers in that they be terminated and restarted without interfering with other Containers. However, some Containers (such as DART) have dependencies on other Containers (FLARE).
In addition to the FLARE and DART Containers, a Global Memory Services (GMS) Container provides memory management functionality and services other Containers. As an example, the FLARE Container takes 500MB memory, while the DART Container takes 2.5GB, all allocated by the GMS.
A kernel space Container is responsible for allocating resources on behalf of user space Containers. The Linux Upstart software provides a means to start, stop and restart Containers.
Control Path
The Control Path is a implemented using technology derived from the Celerra Control Station (itself a Linux-based server) and the CLARiiON NaviSphere software. The Control Path is also where the Common Security Toolkit (CST) is found. The CST appears to be RSA technology and is used in multiple EMC products for security-relation functions. In contrast to the Data Path which consists of functionality directly relating to the transferring of data, the Control Path is concerned with management functionality.
Within the Control Path of the VNXe is the EMC CIM (Common Interface Module) Object Manager (ECOM) management server. ECOM interfaces with "Providers" which are essentially plug-ins. Within the VNXe, ECOM runs on the master SP.
There are a number of different Providers. These include Application Providers for Exchange, iSCSI, Shared Folders and VMware software provision, a Virtual Server Provider, Pools Provider, CLARiiON Provider and Celerra Provider. There are also providers for Registration, Scripting, Scheduling, Replication etc. As plug-ins to the ECOM server, additional services can be written to extend the functionality within the VNXe.
With the use of Providers, ECOM implements a middleware subsystem that can be called by front end applications such as Unisphere or the VNXe command line.
In addition to running Providers, ECOM also provides basic web server functionality used by the Unisphere GUI and CLI via the Apache web server.
Pulling it together
The VNXe uses some additional Linux software along with the custom CSX and ECOM components. High availability is implemented through the open source Pacemaker cluster resource manager and using the Softdog software timing kernel driver. CSX components are resource managed using the cgroups feature of the Linux kernel. The Logging system uses the Postgres database. Although this is covered in the EMC training, it's also possible to see this by checking the output of "ps" from an SSH session.
To understand how the various components hang together, the boot sequence looks a bit like this:
- BIOS/POST
- Linux boots and initiates run level 3
- The "C4" stack is loaded by the Linux Upstart software:
- CSX infra
- Log daemon
- GMS Container
- FLARE Container
- admin
- Pacemaker is loaded and automatically starts:
- Logging
- DART Container
- Control Path software (ECOM on the master SP based on mgmt network status)
Hopefully this gives some insight into the complexity that underpins the VNXe. We're going to look at one more topic to conclude this mini series, and it's a subject that is the source of many questions on the EMC VNXe Community forum. In the next post we'll have a look at VNXe networking...
Wednesday, 4 April 2012
EMC VNXe - diving under the hood (Part 3: DART)
In the previous post, we looked at the parts of the VNXe that are derived from the FLARE (CLARiiON) code. The result is a number of LUNs that are presented up the stack to the DART (Celerra) part of the system.
Using the "svc_storagecheck -l" command, we can see that a total of 20 disks are found. These map to the two FLARE LUNs from the 300GB SAS RAID5 RAID Group and the sixteen FLARE LUNs from the 2TB NL-SAS RAID6 RAID Group, plus two other disks: root_disk and root_ldisk.
root_disk and root_ldisk appear to map to the internal SSD on the Service Processors and are not visible to the end user for configuration. These disks appear to have root filesystems, panic reservation and UFS log filesystems.
The FLARE LUNs are seen as disks to DART and are commonly referred to as "dvols".
The dvols are grouped into Storage Pools. The following are defined by the system, along with a subset of their parameters:
As the above table shows, the LUNs presented from the FLARE side of the VNXe are assigned to the performance_dart0 and capacity_dart1 pools.
The Volume Profile should be familiar to Celerra administrators and is the set of rules that define how a set of disks should be configured.
On a Celerra, disks could be configured manually (if you know exactly what you want) or automatically using the "Automatic Volume Manager" (AVM). Because the VNXe is designed to be simple, AVM does all the work.
An AVM group called "root_avm_vol_group_63" (the svc_neo_map command refers to this as the "Internal FS name") has been created and consists of two dvols, d18 and d19 that corresponds to the performance_dart0 storage pool. These two dvols map to the two LUNs presented from the 300GB SAS disk RAID Group. It appears when a filesystem is created, the first disk is partitioned into a number of slices (sixteen on d18). Each slice then has a volume created on it and finally, another volume is created that spans across all the other volumes. It's this top level volume, called v139 in the diagram below, on which a filesystem is created:
Note that d19 in the above diagram isn't used. If the filesystem is expanded beyond the capacity of the single disk, then presumably the next disk is used. For some reason, slice 68 doesn't have a corresponding volume. I would welcome any explanation as to why this is.
The configuration for the capacity_dart1 pool is very similar, albeit with many more disks (sixteen instead of two) and many more slices. Unfortunately it's too big to show here. As an example, the first disk, d23, has 40 slices of its own that form part of the pool.
The use of all these smaller slices presumably means that a filesystem can grow incrementally from the pool (and possibly shrink?).
When the filesystem is created, it isn't visible to an external host. On a Celerra or VNX, this functionality would be handled by a physical data mover. The VNXe uses a software "Shared Folder Server" (SFS) which acts as the server to the other hosts on the network.
Multiple Shared Folder Servers can be created (apparently up to 12 Shared Folder Servers (file) and/or iSCSI Servers (block) are supported), each with its own network settings and sharing its own filesystems out over NFS or CIFS. Note that while a SFS can handle both NFS and CIFS, a single filesystem within a SFS can support either NFS or CIFS, but not both at the same time.
From a disk perspective, EMC have done well to hide a lot of legacy cruft away from the user and the encapsulation of FLARE and DART, along with the software implementation of the data mover idea is a neat evolution of an aging architecture.
There is more to look into such as networking (which has provoked a significant number of questions on the EMC forums) and I'd like to find out more about the CSX "execution environment" that underpins much of the new design. I'd be sure to post more if/when I get more information, but hopefully you've found this a useful dive under the hood of the VNXe.
Using the "svc_storagecheck -l" command, we can see that a total of 20 disks are found. These map to the two FLARE LUNs from the 300GB SAS RAID5 RAID Group and the sixteen FLARE LUNs from the 2TB NL-SAS RAID6 RAID Group, plus two other disks: root_disk and root_ldisk.
root_disk and root_ldisk appear to map to the internal SSD on the Service Processors and are not visible to the end user for configuration. These disks appear to have root filesystems, panic reservation and UFS log filesystems.
The FLARE LUNs are seen as disks to DART and are commonly referred to as "dvols".
The dvols are grouped into Storage Pools. The following are defined by the system, along with a subset of their parameters:
| Name | Description | In use | Members | Volume Profile |
|---|---|---|---|---|
| clarsas_archive | CLARiiON RAID5 on SAS | False | clarsas_archive_vp | |
| clarsas_r6 | CLARiiON RAID6 on SAS | False | clarsas_r6_vp | |
| clar_r1_3d_sas | 3 disk RAID-1 | False | clar_r1_3d_sas_vp | |
| clar_r3_3P1_SAS | RAID-3 (3+1) | False | clar_r3_3P1_SAS_vp | |
| performance_dart0 | performance | True | d18,d19 | N/A |
| capacity_dart1 | capacity | True | d23,d24,d25,d26 d27,d28,d29,d30 d31,d32,d33,d34 d35,d36,d37,d38 |
N/A |
As the above table shows, the LUNs presented from the FLARE side of the VNXe are assigned to the performance_dart0 and capacity_dart1 pools.
The Volume Profile should be familiar to Celerra administrators and is the set of rules that define how a set of disks should be configured.
On a Celerra, disks could be configured manually (if you know exactly what you want) or automatically using the "Automatic Volume Manager" (AVM). Because the VNXe is designed to be simple, AVM does all the work.
An AVM group called "root_avm_vol_group_63" (the svc_neo_map command refers to this as the "Internal FS name") has been created and consists of two dvols, d18 and d19 that corresponds to the performance_dart0 storage pool. These two dvols map to the two LUNs presented from the 300GB SAS disk RAID Group. It appears when a filesystem is created, the first disk is partitioned into a number of slices (sixteen on d18). Each slice then has a volume created on it and finally, another volume is created that spans across all the other volumes. It's this top level volume, called v139 in the diagram below, on which a filesystem is created:
Note that d19 in the above diagram isn't used. If the filesystem is expanded beyond the capacity of the single disk, then presumably the next disk is used. For some reason, slice 68 doesn't have a corresponding volume. I would welcome any explanation as to why this is.
The configuration for the capacity_dart1 pool is very similar, albeit with many more disks (sixteen instead of two) and many more slices. Unfortunately it's too big to show here. As an example, the first disk, d23, has 40 slices of its own that form part of the pool.
The use of all these smaller slices presumably means that a filesystem can grow incrementally from the pool (and possibly shrink?).
When the filesystem is created, it isn't visible to an external host. On a Celerra or VNX, this functionality would be handled by a physical data mover. The VNXe uses a software "Shared Folder Server" (SFS) which acts as the server to the other hosts on the network.
Multiple Shared Folder Servers can be created (apparently up to 12 Shared Folder Servers (file) and/or iSCSI Servers (block) are supported), each with its own network settings and sharing its own filesystems out over NFS or CIFS. Note that while a SFS can handle both NFS and CIFS, a single filesystem within a SFS can support either NFS or CIFS, but not both at the same time.
From a disk perspective, EMC have done well to hide a lot of legacy cruft away from the user and the encapsulation of FLARE and DART, along with the software implementation of the data mover idea is a neat evolution of an aging architecture.
There is more to look into such as networking (which has provoked a significant number of questions on the EMC forums) and I'd like to find out more about the CSX "execution environment" that underpins much of the new design. I'd be sure to post more if/when I get more information, but hopefully you've found this a useful dive under the hood of the VNXe.
Subscribe to:
Posts (Atom)

