Welcome!

ColdFusion Authors: Yakov Fain, Maureen O'Gara, Nancy Y. Nee, Tad Anderson, Daniel Kaar

Related Topics: ColdFusion

ColdFusion: Article

COSMOS: Managing the ColdFusion Experience

COSMOS: Managing the ColdFusion Experience

Far too often we listen to the naysayers who tell us that something can't be done and give poorly founded reasons as to why our troubles persist. The ColdFusion Application Server is no exception to their folly. If you ask people for the drawbacks of ColdFusion, most will reply "speed" or "stability." Let me be the first to tell you that it does not have to be that way.

The system described in this article was built to change the way we think about our code and applications. COSMOS was designed to change our perceptions of the ColdFusion Application Server and to enhance the ColdFusion experience.

If you have ever looked in the /cfusion/log/ directory you've probably seen one or more of the many ColdFusion-generated error/information logs. These text files can easily grow to hundreds of MB and contain the best indicators of "what happened." As with any other service or application, a regular review of system logs should be a part of normal administration. Unfortunately, because of their large size and the fact that the data is segmented into so many logs, it's difficult to get a complete picture of performance, problems, and failure.

Developers who work on a dedicated server can use the ColdFusion Administrator to view these logs. This can be accomplished by clicking on "Log Files" and then downloading the entire log via a browser. Unfortunately, this is usually not possible given the size of most logs and the remote connection speed.

For shared developers, the critical information is unavailable due to the nature of the shared environment and security. In most cases, developers know only what a site user tells them or what they trap using CFTRY/CFCATCH and CFERROR. Even with these mechanisms in place, the larger picture is unavailable and the majority of performance issues go unnoticed and unattended.

The above issues hinder administrators and developers alike. The result is:

  • No true time or site correlation for ColdFusion Application Server events.
  • Time is wasted attempting to data mine text.
  • Site administrators, developers, and business owners don't know there is a problem.
  • A negative stigma is created based on a lack of timely and organized information.

    To be successful, a solution must have several characteristics:

  • Run autonomously, centrally, and constantly
  • Contain error lookup with "clean code" examples
  • Return all logs and bounced e-mails
  • Bring symmetry to the generated data through trending, aggregation, and normalization
  • Be fast without affecting performance on the managed server

    The solution is COSMOS. Written mainly with Cold-Fusion, it's an integration of ASP, DOS, Perl, ADSI, and Call-XML. It's a remote management platform that leverages the file system, registry, metabase, service controls, and performance counters. Currently, COSMOS contains over 16-million server events aggregated into an MS-SQL database. Captured within a maximum of 40 seconds, these events include all of the following:

    • Application errors
    • CF Application Server stop/starts
    • Hung threads
    • Long-running templates
    • Missing templates
    • Scheduled task results
    • Undeliverable e-mails
    • Mail sent
    How does this affect you? By returning timely, accurate, and relevant information, a developer can see immediately where improvement is needed. At Hostcentric, it's now possible to access several views, graphs, and aggregations that provide a new perspective on the ColdFusion experience.

    There are over 20 COSMOS reports available to a dedicated client, most of which are also available for shared customers. The following is a list of some reports with a brief description of how they impact the development and maintenance cycle.

    Information Listings
    There are several listings available, each with similar characteristics. A listing allows the user to select the maximum number of records to view per screen and how far back to examine data. It also allows the user to progress backward from that point to review previous messages. The majority of listings allow filtering to a single IIS root. They also provide direct access to the complete original error and a corrected code context lookup.

    General Application Error Listing
    Application errors are the best view into the progress and developmental completeness of a site (see Figure 1). A well-coded site generates no application errors. This listing provides a top-down view of the most recent application errors for all IIS roots. By clicking on the error message on the right, a popup window displays the error message as displayed to a site visitor.

    General Missing Template
    This applies to all .cfm templates requested by the Web server, but not found. In most cases, the developer doesn't even know that people are getting "404 File Not Found" messages. If a search engine indexes your site or a user bookmarks a page, a change in the site causes missed business. The solution is to use the default missing template handler in ColdFusion Administrator or to add a CFERROR TYPE="REQUEST" in your site's Application.cfm.

    Long-Running Template Listing
    This applies to the processing time for pages that take longer than expected. The determination of how long is too long is configured in the logging/settings section of ColdFusion Administrator. A typical setting is 45 seconds, though anything taking that long would most likely be canceled or ignored by the calling client. In addition, a script running for 45 seconds could help identify a performance bottleneck for the application server. By default, CF Administrator doesn't enable this counter. Over time, a development team should ratchet this value as low as possible to get the best diagnostics.

    Undeliverable CFMAIL Listing
    When ColdFusion is unable to deliver a message to the server specified in a CFMAIL script, the original template is renamed and filed in the /cfusion/ mail/undeliver/ directory. An error message is also written to the Mail.log or Error.log describing the problem that prevents proper delivery. This listing binds those two pieces of information together.

    The following popup allows an administrator to correct and resend the message from the original server. This function is indispensable for any business that relies on CFMAIL to reliably carry e-mail, and can't accept undelivered messages.

    Hung Thread Listing
    This is probably the greatest indicator of a performance and stability problem. Hung threads are ColdFusion's method of alerting us that it was unable to completely process the requested template. This is usually the result of code or database issues. CF4.x and above has an option in the Administrator to have CF "restart at x unresponsive requests."

    When the hung thread count matches the defined threshold, ColdFusion reaches a critical point and will stop/restart itself to avoid excessive downtime. Constant examination of hung threads is necessary to avoid application server failure. At the end of this article I've included three links that help to define more fully the causes of hung threads.

    Scheduled Task Listing
    Most scheduled tasks run completely unnoticed until someone realizes that a critical function has not processed in days. This listing is not much to look at but, under the hood, a huge modification and improvement has been created for the executive service.

    As always, COSMOS can determine if your task started, succeeded, or failed based on the logs. Furthermore, COSMOS will allow you to define a target string in the page HTML and record the generated content from the target URL to the database. If a scheduled task does not return the defined string, an e-mail containing the content and diagnostics can be generated at the time of failure. In addition, the actual HTTP response (CFHTTP.FILECONTENT) is zipped and written to the database.

    Aggregation and Stratification
    More commonly called a GROUPING, the next series of graphs were created to help identify the greatest problems quickly. By examining the data based on time, date, and IIS root, we can gather a greater understanding of where faults exist.

    Application Log Stratification by IIS Root
    Over a selectable time span, this graph allows you to see which sites are having the greatest incidence of errors (see Figure 2). By clicking on the blue horizontal bar on the right, you're driven back to the general application error listing but with an additional sort parameter that isolates errors created by the target root.

    Time/Error Graph
    Especially useful in determining if your day is getting better or worse, this graph breaks down the server errors by 10 minute increments over a selectable date span. This is often used to diagnose a recurring failure point over a multiple day or week period.

    Application Errors Stratified by Date
    Similar to the previous idea, this graph groups the number of errors by the date that they occurred (see Figure 3). This helps to identify programming trends and can easily indicate a "bad day" for an application. By clicking on the blue bar, your browser is taken to the application log stratification by IIS root. Clicking on the "Time Graph" button brings you to the next graph.

    Long-Running Template Aggregation by IIS Root
    Similar to the previous root aggregations, this has several prominent exceptions. Because a long-running page has a value associated with the processing time, I've included a column for the sum and average values. Using this display, it's possible to extract the templates most often run beyond acceptable limits, thus demanding the greatest processing time. This affects performance, though not necessarily a failure, and is a fantastic indicator of templates that need to be addressed before they become a stability issue.

    Hung Thread Aggregation by IIS Root
    This graph will often tell which application is responsible for killing the server. Over a selectable data span you can easily see which sites are causing CF to lose processing threads and tie up resources. The blue horizontal bar links back to the hung thread listing for a given root.

    The Big and the Bad
    With the thousands of errors, tasks, and events returned each hour, it's easy to become overwhelmed. In addition, not all errors have the same weight on the application server or urgency to a business owner. To resolve this problem a system of alerts and probes runs in the background. Operating at several intervals, the most relevant problems are quickly pulled out of the pool and matched with a type and severity. Once a probe identifies an event or error candidate, the alert has an option of paging, e-mailing, or calling (using CallXML) the administrator.

    A Final Look
    When did your application server last crash and why?

  • Event Chronology: As the first view that brings together data from multiple sources (see Figure 4), it provides a chronological view of all application errors, hung threads, long-running templates, and application server failures. The graph threads events are based on time in order to provide a trace leading up to a failure.
  • Spectral Analysis: This graph is unique because it rapidly identifies problems that would otherwise slip under the wire (see Figure 5). The three colors representing CF stops (red), starts (green), and hung threads (purple) are graphed relative to a 24-hour time line. By viewing all hung threads that led to server failures, a complete understanding of the root performance issues is garnered.

    Summary
    Tonight at 1 a.m. your database is going to run out of space and begin throwing application errors. Maybe your mail server stops relaying your order confirmations. There are a thousand permutations to a preventable and containable failure. Will your customers be the first to let you know?

    In the end, owners and developers, shared and dedicated, all have the same concerns: stability and performance. Using this system makes that realization no more than a few seconds away.

    Related Articles

  • http://allaire.com/Handlers/index.cfm?ID=8627&Method=Full
  • http://allaire.com/Handlers/index.cfm?ID=2497&Method=Full
  • http://allaire.com/Handlers/index.cfm?ID=1505&Method=Full
  • http://allaire.com/Handlers/index.cfm?ID=1540&Method=Full
  • http://support.microsoft.com/support/kb/articles/q174/4/96.asp
  • More Stories By Tim Nettleton

    Tim Nettleton is a senior engineer at Hostcentric’s Orlando division. He has worked with ColdFusion for four years and recently spoke at DevCon 2001 in Orlando, Florida.

    Comments (5) View Comments

    Share your thoughts on this story.

    Add your comment
    You must be signed in to add a comment. Sign-in | Register

    In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


    Most Recent Comments
    Craig Rosenblum 05/29/03 02:24:00 PM EDT

    I am not interested in the whole system mainly right now, the log component.

    It is a good idea, to help manage logs, and make sure we are really on top of errors.

    And it does look nice as well.

    jurgen koch 03/11/02 04:48:00 PM EST

    ok...let me retract the attitude from my immediately previous post and clarify the situation with more fact that flame.

    I contacted the articles author, Timothy Nettleton ([email protected]), who was very courteous and replied almost immediately to explain that while Cosmos will probably not be released for sale by Hostcentric, they can provide you remote access to Cosmos through its web interface at cosmos.hostcentric.net

    You will have to work out details re: how Hostcentric will obtain access to your CF log files for parsing, etc. but the good news is that Cosmos is available to the public...and the pricing is very reasonable.

    embarrased by my previous rant,
    Jurgen

    jurgen koch 03/11/02 01:19:00 PM EST

    I contacted Hostcentric (the ISP that the author of the article worked for) and the impression that I got was that
    COSMOS is only available to their colocation, etc. clients, and while they did develop the application, they have no plans to release it to the public.

    Kinda makes me wonder why CFDJ published the article at all. Just to tease CF Administrators?

    Hey CFDJ, want to pay me to write an article about the really cool CF apps I have developed, but can't/won't sell, or release code, logic, or the application to the public? Thanks for nothing.

    Greg Correll 02/15/02 12:55:00 PM EST

    Is cosmos available to the public?

    Matt McDonald 01/17/02 12:16:00 AM EST

    Is cosmos available to the public?

    @ThingsExpo Stories
    SYS-CON Events announced today that Matrix.org has been named “Silver Sponsor” of Internet of @ThingsExpo, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Matrix is an ambitious new open standard for open, distributed, real-time communication over IP. It defines a new approach for interoperable Instant Messaging and VoIP based on pragmatic HTTP APIs and WebRTC, and provides open source reference implementations to showcase and bootstrap the new standard. Our focus is on simplicity, security, and supporting the fullest feature set.
    WebRTC defines no default signaling protocol, causing fragmentation between WebRTC silos. SIP and XMPP provide possibilities, but come with considerable complexity and are not designed for use in a web environment. In his session at Internet of @ThingsExpo, Matthew Hodgson, technical co-founder of the Matrix.org, will discuss how Matrix is a new non-profit Open Source Project that defines both a new HTTP-based standard for VoIP & IM signaling and provides reference implementations.

    SUNNYVALE, Calif., Oct. 20, 2014 /PRNewswire/ -- Spansion Inc. (NYSE: CODE), a global leader in embedded systems, today added 96 new products to the Spansion® FM4 Family of flexible microcontrollers (MCUs). Based on the ARM® Cortex®-M4F core, the new MCUs boast a 200 MHz operating frequency and support a diverse set of on-chip peripherals for enhanced human machine interfaces (HMIs) and machine-to-machine (M2M) communications. The rich set of periphera...

    SYS-CON Events announced today that Aria Systems, the recurring revenue expert, has been named "Bronze Sponsor" of SYS-CON's 15th International Cloud Expo®, which will take place on November 4-6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Aria Systems helps leading businesses connect their customers with the products and services they love. Industry leaders like Pitney Bowes, Experian, AAA NCNU, VMware, HootSuite and many others choose Aria to power their recurring revenue business and deliver exceptional experiences to their customers.
    The Internet of Things (IoT) is going to require a new way of thinking and of developing software for speed, security and innovation. This requires IT leaders to balance business as usual while anticipating for the next market and technology trends. Cloud provides the right IT asset portfolio to help today’s IT leaders manage the old and prepare for the new. Today the cloud conversation is evolving from private and public to hybrid. This session will provide use cases and insights to reinforce the value of the network in helping organizations to maximize their company’s cloud experience.
    The Internet of Things (IoT) is making everything it touches smarter – smart devices, smart cars and smart cities. And lucky us, we’re just beginning to reap the benefits as we work toward a networked society. However, this technology-driven innovation is impacting more than just individuals. The IoT has an environmental impact as well, which brings us to the theme of this month’s #IoTuesday Twitter chat. The ability to remove inefficiencies through connected objects is driving change throughout every sector, including waste management. BigBelly Solar, located just outside of Boston, is trans...
    SYS-CON Events announced today that Matrix.org has been named “Silver Sponsor” of Internet of @ThingsExpo, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Matrix is an ambitious new open standard for open, distributed, real-time communication over IP. It defines a new approach for interoperable Instant Messaging and VoIP based on pragmatic HTTP APIs and WebRTC, and provides open source reference implementations to showcase and bootstrap the new standard. Our focus is on simplicity, security, and supporting the fullest feature set.
    Predicted by Gartner to add $1.9 trillion to the global economy by 2020, the Internet of Everything (IoE) is based on the idea that devices, systems and services will connect in simple, transparent ways, enabling seamless interactions among devices across brands and sectors. As this vision unfolds, it is clear that no single company can accomplish the level of interoperability required to support the horizontal aspects of the IoE. The AllSeen Alliance, announced in December 2013, was formed with the goal to advance IoE adoption and innovation in the connected home, healthcare, education, aut...
    SYS-CON Events announced today that Red Hat, the world's leading provider of open source solutions, will exhibit at Internet of @ThingsExpo, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Red Hat is the world's leading provider of open source software solutions, using a community-powered approach to reliable and high-performing cloud, Linux, middleware, storage and virtualization technologies. Red Hat also offers award-winning support, training, and consulting services. As the connective hub in a global network of enterprises, partners, a...
    The only place to be June 9-11 is Cloud Expo & @ThingsExpo 2015 East at the Javits Center in New York City. Join us there as delegates from all over the world come to listen to and engage with speakers & sponsors from the leading Cloud Computing, IoT & Big Data companies. Cloud Expo & @ThingsExpo are the leading events covering the booming market of Cloud Computing, IoT & Big Data for the enterprise. Speakers from all over the world will be hand-picked for their ability to explore the economic strategies that utility/cloud computing provides. Whether public, private, or in a hybrid form, clo...
    Software AG helps organizations transform into Digital Enterprises, so they can differentiate from competitors and better engage customers, partners and employees. Using the Software AG Suite, companies can close the gap between business and IT to create digital systems of differentiation that drive front-line agility. We offer four on-ramps to the Digital Enterprise: alignment through collaborative process analysis; transformation through portfolio management; agility through process automation and integration; and visibility through intelligent business operations and big data.
    The Transparent Cloud-computing Consortium (abbreviation: T-Cloud Consortium) will conduct research activities into changes in the computing model as a result of collaboration between "device" and "cloud" and the creation of new value and markets through organic data processing High speed and high quality networks, and dramatic improvements in computer processing capabilities, have greatly changed the nature of applications and made the storing and processing of data on the network commonplace.
    Be Among the First 100 to Attend & Receive a Smart Beacon. The Physical Web is an open web project within the Chrome team at Google. Scott Jenson leads a team that is working to leverage the scalability and openness of the web to talk to smart devices. The Physical Web uses bluetooth low energy beacons to broadcast an URL wirelessly using an open protocol. Nearby devices can find all URLs in the room, rank them and let the user pick one from a list. Each device is, in effect, a gateway to a web page. This unlocks entirely new use cases so devices can offer tiny bits of information or simple i...
    Things are being built upon cloud foundations to transform organizations. This CEO Power Panel at 15th Cloud Expo, moderated by Roger Strukhoff, Cloud Expo and @ThingsExpo conference chair, will address the big issues involving these technologies and, more important, the results they will achieve. How important are public, private, and hybrid cloud to the enterprise? How does one define Big Data? And how is the IoT tying all this together?
    The Internet of Things (IoT) is going to require a new way of thinking and of developing software for speed, security and innovation. This requires IT leaders to balance business as usual while anticipating for the next market and technology trends. Cloud provides the right IT asset portfolio to help today’s IT leaders manage the old and prepare for the new. Today the cloud conversation is evolving from private and public to hybrid. This session will provide use cases and insights to reinforce the value of the network in helping organizations to maximize their company’s cloud experience.
    TechCrunch reported that "Berlin-based relayr, maker of the WunderBar, an Internet of Things (IoT) hardware dev kit which resembles a chunky chocolate bar, has closed a $2.3 million seed round, from unnamed U.S. and Switzerland-based investors. The startup had previously raised a €250,000 friend and family round, and had been on track to close a €500,000 seed earlier this year — but received a higher funding offer from a different set of investors, which is the $2.3M round it’s reporting."
    The Industrial Internet revolution is now underway, enabled by connected machines and billions of devices that communicate and collaborate. The massive amounts of Big Data requiring real-time analysis is flooding legacy IT systems and giving way to cloud environments that can handle the unpredictable workloads. Yet many barriers remain until we can fully realize the opportunities and benefits from the convergence of machines and devices with Big Data and the cloud, including interoperability, data security and privacy.
    All major researchers estimate there will be tens of billions devices - computers, smartphones, tablets, and sensors - connected to the Internet by 2020. This number will continue to grow at a rapid pace for the next several decades. Over the summer Gartner released its much anticipated annual Hype Cycle report and the big news is that Internet of Things has now replaced Big Data as the most hyped technology. Indeed, we're hearing more and more about this fascinating new technological paradigm. Every other IT news item seems to be about IoT and its implications on the future of digital busines...
    Cultural, regulatory, environmental, political and economic (CREPE) conditions over the past decade are creating cross-industry solution spaces that require processes and technologies from both the Internet of Things (IoT), and Data Management and Analytics (DMA). These solution spaces are evolving into Sensor Analytics Ecosystems (SAE) that represent significant new opportunities for organizations of all types. Public Utilities throughout the world, providing electricity, natural gas and water, are pursuing SmartGrid initiatives that represent one of the more mature examples of SAE. We have s...
    The Internet of Things needs an entirely new security model, or does it? Can we save some old and tested controls for the latest emerging and different technology environments? In his session at Internet of @ThingsExpo, Davi Ottenheimer, EMC Senior Director of Trust, will review hands-on lessons with IoT devices and reveal privacy options and a new risk balance you might not expect.